Mmcp.market

api-rate-limit-handler skill

by sickn33·sickn33/agentic-awesome-skills·47k stars·MIT

Implement bounded, idempotency-aware API throttling, backoff, and retry handling for 429 and transient 5xx responses.

A100/100content scan

Is the api-rate-limit-handler skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the api-rate-limit-handler skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git /tmp/agentic-awesome-skills
mkdir -p ~/.claude/skills
cp -r /tmp/agentic-awesome-skills/plugins/agentic-awesome-skills-claude/skills/api-rate-limit-handler ~/.claude/skills/api-rate-limit-handler
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

API Rate Limit Handler

Overview

A skill for implementing production-grade rate limiting, exponential backoff, and retry strategies when integrating with external APIs. Prevents cascading failures, respects upstream quotas, and keeps your application resilient under load.

When to Use This Skill

  • Use when calling external APIs that enforce rate limits (OpenAI, Stripe, GitHub, etc.)
  • Use when you receive 429 Too Many Requests or 5xx errors and need graceful recovery
  • Use when building a client that must respect Retry-After headers
  • Use when designing a system that fans out to multiple API providers
  • Use when the user says "handle rate limits", "add retry logic", "backoff strategy", or "don't get throttled"

How It Works

Step 1: Classify the response

Determine whether a failed request is retryable or terminal.

Step 2: Parse rate limit headers

Always check upstream hints before computing your own delay.

function getRetryDelay(
  response: Response,
  attempt: number,
  maxDelayMs = 60_000
): number {
  // Prefer upstream hints
  const retryAfter = response.headers.get("Retry-After");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds) && seconds >= 0) {
      return Math.min(seconds * 1000, maxDelayMs);
    }
    // HTTP-date format
    const date = new Date(retryAfter).getTime();
    if (Number.isFinite(date)) {
      return Math.min(Math.max(0, date - Date.now()), maxDelayMs);
    }
  }

  // GitHub documents x-ratelimit-reset as Unix epoch seconds.
  const githubReset = Number(response.headers.get("x-ratelimit-reset"));
  if (Number.isFinite(githubReset)) {
    return Math.min(
      Math.max(0, githubReset * 1000 - Date.now()),
      maxDelayMs
    );
  }

  // Fallback: capped exponential backoff with full jitter.
  const cap = Math.min(1000 * 2 ** attempt, maxDelayMs);
  return Math.floor(Math.random() * cap);
}

Provider-specific reset headers do not share one unit or format. For example, some APIs return durations while GitHub returns epoch seconds. Parse an additional header only after checking that provider's current documentation.

Step 3: Implement the retry loop

async function fetchWithRetry(
  url: string,
  options: RequestInit,
  maxRetries = 3,
  maxElapsedMs = 120_000,
  retryNonIdempotent = false
): Promise<Response> {
  const startedAt = Date.now();
  const method = (options.method ?? "GET").toUpperCase();
  const replaySafe = ["GET", "HEAD", "OPTIONS", "PUT", "DELETE"].includes(method)
    || retryNonIdempotent;

  for (let attempt = 0; attempt <= maxRetries; attempt++) {
    const response = await fetch(url, options);

    if (response.ok) return response;

    // Terminal errors — do not retry
    if ([400, 401, 403, 404, 422].includes(response.status)) {
      throw new Error(`Terminal error ${response.status}: ${response.statusText}`);
    }

    if (!replaySafe) {
      throw new Error(
        `${method} was not retried because replay safety was not explicitly established`
      );
    }

    // Retryable — but exhausted attempts
    if (attempt === maxRetries) {
      throw new Error(`Failed after ${maxRetries} retries: ${response.status}`);
    }

    const remaining = maxElapsedMs - (Date.now() - startedAt);
    const delay = Math.min(getRetryDelay(response, attempt), remaining);
    if (delay <= 0) {
      throw new Error

Step 4: Add a client-side rate limiter (proactive)

Prevent hitting upstream limits in the first place with a token bucket or sliding window.

class TokenBucket {
  private tokens: number;
  private lastRefill: number;
  private queue: Promise<void> = Promise.resolve();

  constructor(
    private maxTokens: number,
    private refillRate: number // tokens per second
  ) {
    this.tokens = maxTokens;
    this.lastRefill = Date.now();
  }

  async acquire(): Promise<void> {
    const ticket = this.queue.then(() => this.acquireOnce());
    this.queue = ticket.catch(() => undefined);
    return ticket;
  }

  private async acquireOnce(): Promise<void> {
    this.refill();
    if (this.tokens < 1) {
      const waitMs = ((1 - this.tokens) / this.refillRate) * 1000;
      await new Promise(resolve => setTimeout(resolve, waitMs));
      this.refill();
    }
    this.tokens -= 1;
  }

  private refill(): void {
    const now = Date.now();
    const elapsed = (now - this.lastRefill) / 1000;
    this.tokens = Math.min(this.maxTokens, this.tokens + elapsed * this.refillRate);
    this.lastRefill = now;
  }
}

// Usage: limit to 60 requests/minute
const limiter = new TokenBucket(60, 1);

async function rateLimitedFetch(url: string, options: RequestInit) {
  await limiter.acquire();
  return fetchWithRetry(url, options);
}

Examples

Example 1: Idempotent API read with retry

const response = await fetchWithRetry(
  "https://api.github.com/repos/OWNER/REPO",
  {
    method: "GET",
    headers: {
      "Accept": "application/vnd.github+json",
      "Authorization": `Bearer ${githubToken}`,
    },
  },
  3
);

For a POST or another operation with side effects, leave retryNonIdempotent false unless the provider documents an idempotency mechanism and the same stable idempotency key is reused for every attempt.

Example 2: Python implementation

import time
import random
import httpx

def fetch_with_retry(url: str, max_retries: int = 3, **kwargs) -> httpx.Response:
    for attempt in range(max_retries + 1):
        response = httpx.request("GET", url, **kwargs)

        if response.is_success:
            return response

        if response.status_code in (400, 401, 403, 404, 422):
            response.raise_for_status()

        if attempt == max_retries:
            response.raise_for_status()

        # Parse Retry-After or compute backoff
        retry_after = response.headers.get("retry-after")
        if retry_after and retry_after.isdigit():
            delay = int(retry_after)
        else:
            delay = min(2 ** attempt + random.uniform(0, 1), 60)

        print(f"Retrying in {delay:.1f}s (attempt {attempt + 1}/{max_retries})")
        time.sleep(delay)

    raise RuntimeError("Unreachable")

Best Practices

  • ✅ Always respect Retry-After headers — they come from the provider who knows their limits
  • ✅ Add jitter to backoff to prevent thundering herd when multiple clients retry simultaneously
  • ✅ Log every retry with status code, delay, and attempt number for debugging
  • ✅ Set a maximum total timeout to avoid hanging indefinitely
  • ✅ Use a client-side rate limiter proactively rather than only reacting to 429s
  • ✅ Retry state-changing requests only with a provider-documented idempotency mechanism and a stable key
  • ❌ Don't retry 4xx client errors (except 408 and 429) — fix the request instead
  • ❌ Don't use fixed delays — exponential backoff distributes load more evenly
  • ❌ Don't retry without a cap — unbounded retries can amplify outages
  • ❌ Don't ignore per-endpoint limits — some APIs have different quotas per route

Limitations

  • This skill does not replace environment-specific validation, testing, or expert review.
  • Token bucket is approximate for distributed systems — use Redis-backed rate limiting for multi-instance deployments (for example the upstash-ratelimit skill, or any shared-store limiter).
  • Some APIs use non-standard rate limit headers; check provider documentation.
  • The elapsed-time cap shown here bounds retry waits, not a single hung network call; combine it with an AbortSignal or client timeout.

Common Pitfalls

Solution: Use exponential backoff with jitter and a circuit breaker for sustained failures.

  • Problem: Retrying too aggressively during an outage amplifies the problem.

Solution: Add randomized jitter (Math.random() 0.3 delay) to decorrelate retries.

  • Problem: Multiple instances of your app all retry at the same time (thundering herd).

Solution: Parse both formats — check if the value is numeric first, then try Date parsing.

  • Problem: Retry-After header contains an HTTP-date instead of seconds.

Solution: Serialize acquisition within one process, decrement before send, and use a shared distributed limiter across instances.

  • Problem: Client-side limiter doesn't account for concurrent requests already in-flight.

Related Skills

  • @poka-yoke - Mistake-proofing APIs so invalid requests never reach the retry path
  • @circuit-breaker - When to stop retrying entirely and fail fast

More skills from sickn33/agentic-awesome-skills

  • A00-andruia-consultantArquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza la hoja de ruta óptima para proyectos de IA en español.
  • F007Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code review, incident response, and infrastructure security for any project.
  • A10-andruia-skill-smithIngeniero de Sistemas de Andru.ia. Diseña, redacta y despliega nuevas habilidades (skills) dentro del repositorio siguiendo el Estándar de Diamante.
  • A20-andruia-niche-intelligenceEstratega de Inteligencia de Dominio de Andru.ia. Analiza el nicho específico de un proyecto para inyectar conocimientos, regulaciones y estándares únicos del sector. Actívalo tras definir el nicho.
  • A2slides-ppt-generatorAI-powered presentation generation via the 2slides API — create slides from text, match a reference image style, summarize documents into decks, add AI voice narration, and export pages/audio. Use for any \"make slides\", \"create a deck\", or \"slides from this document\" request.
  • A3d-web-experienceExpert in building 3D experiences for the web - Three.js, React
  • Aab-test-setupUse when designing an A/B or split test: define the hypothesis, control and variants, estimate sample size, verify tracking, and predeclare metrics and stopping rules.
  • Aab-testingWhen the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
  • Aacceptance-orchestratorUse when a coding task should be driven end-to-end from issue intake through implementation, review, deployment, and acceptance verification with minimal human re-intervention.
  • Aaccess-reviewConduct periodic access reviews and certifications. Implement access
  • Aaccessibility-compliance-accessibility-auditYou are an accessibility expert specializing in WCAG compliance, inclusive design, and assistive technology compatibility. Conduct audits, identify barriers, and provide remediation guidance.
  • Aaccesslint-auditFind and fix WCAG 2.2 accessibility issues. Two modes — report (sweep a codebase or page, produce a prioritized written report, no edits) and fix (audit→edit→verify loop on a target). Prefers direct-CDP live-DOM auditing; falls back to a browser-MCP composition or HTML-string audits.

All agent skills → · MCP servers