Skip to main content

API reference

Rate limits

Per-key limits for each endpoint group, headers and backoff.

How limits are applied

Rate limits count requests per API key. Each key has its own counters, so two keys in the same organization don't share a limit, and one busy key can't slow another down. IP addresses play no part.

Each limit counts requests in a fixed 60-second window. The window starts with the first request and resets 60 seconds later. It doesn't slide forward with each request.

Each endpoint group has its own counter. If you use up the call-placement limit, your reads aren't affected, and the reverse is also true.

Finn checks the API key before the rate limit. A request rejected with 401 doesn't use up your limit. Neither does a body that isn't valid JSON: it's rejected with 400 before the key is checked. A request with valid JSON that fails validation does use up your limit, because validation happens after the rate-limit check.

The number of calls running at the same time is a separate limit and isn't covered here.

Limits

All limits are per key, per 60 seconds. Endpoints listed in the same row share one counter.

EndpointsLimit
POST /calls60
GET /calls/{call_uuid}, GET /finns, GET /finns/{finn_id}, GET /wallet300
POST /finns, PATCH /finns/{finn_id}, DELETE /finns/{finn_id}60
GET /voices, GET /voices/{voice_id}60
POST /deployments, POST /deployments/inbound10
POST /deployments/{id}/stop, POST /deployments/inbound/{finn_id}/stop30
GET /deployments, GET /deployments/{id}, GET /wallet/transactions300
GET /phone_numbers, GET /phone_numbers/{id}300
GET /phone_numbers/available30
POST /phone_numbers, DELETE /phone_numbers/{id}10
GET /audiences, GET /audiences/{audience_id}, GET /audiences/{audience_id}/contacts300
POST /audiences, DELETE /audiences/{audience_id}, POST and DELETE on /audiences/{audience_id}/contacts60

These limits are the same for every organization. Finn can change them, so your client should follow the response headers rather than hardcoding these numbers.

Response headers

Every response from a rate-limited endpoint includes the current state of its counter, including successful responses:

HeaderMeaning
RateLimit-PolicyThe limit and window, for example 300;w=60
RateLimit-LimitRequests allowed in the current window for this endpoint group
RateLimit-RemainingRequests left in the current window
RateLimit-ResetSeconds until the window resets. This is a countdown, not a Unix timestamp.
Retry-AfterSeconds to wait before retrying. Sent only on 429.

The headers have no X- prefix.

curl -i https://api.hirefinn.ai/api/v1/finns \
  -H "Authorization: Bearer $FINN_API_KEY"
HTTP/1.1 200 OK
RateLimit-Policy: 300;w=60
RateLimit-Limit: 300
RateLimit-Remaining: 297
RateLimit-Reset: 42

When you exceed the limit

You get 429 Too Many Requests with this body:

{
  "error": "rate_limited",
  "message": "Too many requests. Limit is 300 per minute per API key."
}

message names the limit you hit, but don't parse it. Use the headers.

A 429 means Finn rejected the request before doing anything. No call was placed, nothing was written and your wallet wasn't charged, so retrying is always safe.

Backing off

When Retry-After is present, wait that many seconds. If it's missing, use exponential backoff with jitter. If several workers all retry on the same fixed interval, they hit the limit together again.

import os, random, time
import requests

BASE = "https://api.hirefinn.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['FINN_API_KEY']}"}

def get_with_backoff(path, max_attempts=6):
    for attempt in range(max_attempts):
        response = requests.get(f"{BASE}{path}", headers=HEADERS, timeout=30)
        if response.status_code != 429:
            response.raise_for_status()
            return response.json()
        wait = response.headers.get("Retry-After")
        delay = float(wait) if wait is not None else min(2 ** attempt, 60) + random.uniform(0, 1)
        time.sleep(delay)
    raise RuntimeError(f"Rate limited after {max_attempts} attempts: {path}")

print(get_with_backoff("/finns"))

This TypeScript version slows down before it reaches the limit, so it rarely gets a 429:

const BASE = "https://api.hirefinn.ai/api/v1";
const KEY = process.env.FINN_API_KEY!;

const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

export async function request(path: string, init: RequestInit = {}) {
  for (let attempt = 0; attempt < 6; attempt++) {
    const response = await fetch(`${BASE}${path}`, {
      ...init,
      headers: { ...init.headers, Authorization: `Bearer ${KEY}` },
    });

    if (response.status === 429) {
      const retryAfter = response.headers.get("Retry-After");
      const delay = retryAfter
        ? Number(retryAfter) * 1000
        : Math.min(2 ** attempt * 1000, 60_000) + Math.random() * 1000;
      await sleep(delay);
      continue;
    }

    const remaining = Number(response.headers.get("RateLimit-Remaining"));
    const reset = Number(response.headers.get("RateLimit-Reset"));
    if (!Number.isNaN(remaining) && remaining < 5 && !Number.isNaN(reset)) {
      await sleep(reset * 1000);
    }

    if (!response.ok) {
      throw new Error(`${response.status} ${await response.text()}`);
    }
    return response.json();
  }
  throw new Error(`Rate limited after 6 attempts: ${path}`);
}

Only retry a write automatically after a 429. After a 5xx or a timeout, follow api-idempotency.

Common mistakes

Small pages. Small page sizes use up the read limit quickly. Fetching 10,000 audience contacts 100 at a time takes 100 requests. Ask for the largest page each endpoint allows. See api-pagination.

Polling for call results. If you check a five-minute call every second, you spend 300 requests on that one call, which is your whole read limit for a minute. Subscribe to call.completed instead, and poll GET /calls/{call_uuid} only to fill in events your receiver missed.

Launching campaigns too quickly. POST /deployments allows 10 per minute per key. Spread out bulk launches, or give a separate job its own key.

Load testing against production. There's no sandbox or test allocation, so a load test uses the same limits as your live traffic and places real calls.

Requesting an increase

The limits above are the same for everyone. If you need more, contact Finn support with your organization ID, the endpoints you need more for, and your expected peak requests per minute.