Goal: a local table of every Data Engineer and Analytics Engineer job in five EU countries, filled once, kept current every day, and repaired after downtime.
Four patterns, one key: the canonical job id. It is stable across sources and requests, so every pattern ends in an upsert on id.
| Pattern | Call | When |
|---|---|---|
| Initial backfill | POST /v1/searches (async) |
Once, up to 10,000 jobs per search |
| Daily incremental | POST /v1/jobs/search with posted_within_days |
Every day |
| Lifecycle updates | Watch with job.closed, job.updated, job.reposted |
Continuous, free |
| Outage recovery | GET /v1/events?since= |
After your side was down |
Your table
Section titled “Your table”Store the full job as JSON plus the columns you query. Guard the upsert so older data never overwrites newer data.
CREATE TABLE jobs ( id text PRIMARY KEY, -- canonical job id, job_... status text NOT NULL, -- open | closed closed_reason text, -- filled | expired | removed | unknown | NULL last_seen_at timestamptz NOT NULL, doc jsonb NOT NULL -- the full Job object);
CREATE TABLE sync_state (key text PRIMARY KEY, value text NOT NULL);
-- Upsert one job from a search page or a job.opened / job.updated eventINSERT INTO jobs (id, status, closed_reason, last_seen_at, doc)VALUES ($1, $2, $3, $4, $5)ON CONFLICT (id) DO UPDATESET status = EXCLUDED.status, closed_reason = EXCLUDED.closed_reason, last_seen_at = EXCLUDED.last_seen_at, doc = EXCLUDED.docWHERE jobs.last_seen_at <= EXCLUDED.last_seen_at;Initial backfill
Section titled “Initial backfill”-
Estimate. Send the filters to
POST /v1/jobs/searchwithdry_run: true. Free. Ifexpected_unique_jobs_range.maxis above 10,000, split the backfill (see below). -
Create an async search.
POST /v1/searcheswithlimitup to 10,000,waterfall.max_creditsas a hard cap, and anIdempotency-Key. You get202and asrch_...id. -
Wait for it. Pass
webhook_urlto receivesearch.completed, or pollGET /v1/searches/{id}. Branch onstatus, not on the HTTP code. -
Read every page. Once
statusiscompletedorpartial,dataholds the first page. Passnext_cursorascursoruntil it isnull. Reading is free. -
Upsert each job on
id.
# 1. Create the search (charged as jobs are collected)curl https://api.betterjobs.cc/v1/searches \ -H "Authorization: Bearer $BETTERJOBS_API_KEY" \ -H "BetterJobs-Version: 2026-10-01" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: 0f6e2a94-7b3c-4d81-9e5a-c2b8d4f61a37" \ -d '{ "filters": { "title_or": ["Data Engineer", "Analytics Engineer"], "country_code_or": ["DE", "FR", "NL", "ES", "PL"], "posted_within_days": 30 }, "waterfall": { "strategy": "max_coverage", "max_credits": 5000 }, "limit": 5000, "webhook_url": "https://hooks.northwind.example/betterjobs" }'
# 2. Check status / read pages (free). Add &cursor=<next_cursor> for the next page.curl "https://api.betterjobs.cc/v1/searches/srch_2Vd9KqL4mN?limit=100" \ -H "Authorization: Bearer $BETTERJOBS_API_KEY" \ -H "BetterJobs-Version: 2026-10-01"import osimport timeimport uuid
import requests
API = "https://api.betterjobs.cc/v1"HEADERS = { "Authorization": f"Bearer {os.environ['BETTERJOBS_API_KEY']}", "BetterJobs-Version": "2026-10-01",}FILTERS = { "title_or": ["Data Engineer", "Analytics Engineer"], "country_code_or": ["DE", "FR", "NL", "ES", "PL"],}
def backfill(days: int, max_credits: int) -> None: created = requests.post( f"{API}/searches", headers={**HEADERS, "Idempotency-Key": str(uuid.uuid4())}, json={ "filters": {**FILTERS, "posted_within_days": days}, "waterfall": {"strategy": "max_coverage", "max_credits": max_credits}, "limit": max_credits, }, timeout=30, ) created.raise_for_status() search_id = created.json()["id"]
while True: page = get_search(search_id) if page["status"] in ("completed", "partial"): break if page["status"] == "failed": raise RuntimeError(f"search {search_id} failed") if page["status"] == "on_hold": print("Out of credits. Top up and the search resumes.") time.sleep(15)
while True: for job in page["data"]: upsert_job(job) # the SQL upsert above if page["next_cursor"] is None: return page = get_search(search_id, page["next_cursor"])
def get_search(search_id: str, cursor: str | None = None) -> dict: params = {"limit": 100, **({"cursor": cursor} if cursor else {})} resp = requests.get(f"{API}/searches/{search_id}", headers=HEADERS, params=params, timeout=30) resp.raise_for_status() return resp.json()
backfill(days=30, max_credits=5000)const API = 'https://api.betterjobs.cc/v1';const headers = { Authorization: `Bearer ${process.env.BETTERJOBS_API_KEY}`, 'BetterJobs-Version': '2026-10-01', 'Content-Type': 'application/json',};const filters = { title_or: ['Data Engineer', 'Analytics Engineer'], country_code_or: ['DE', 'FR', 'NL', 'ES', 'PL'],};const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
async function getSearch(id: string, cursor?: string) { const qs = new URLSearchParams({ limit: '100', ...(cursor ? { cursor } : {}) }); const res = await fetch(`${API}/searches/${id}?${qs}`, { headers }); if (!res.ok) throw new Error(`${res.status} ${await res.text()}`); return res.json();}
async function backfill(days: number, maxCredits: number) { const res = await fetch(`${API}/searches`, { method: 'POST', headers: { ...headers, 'Idempotency-Key': crypto.randomUUID() }, body: JSON.stringify({ filters: { ...filters, posted_within_days: days }, waterfall: { strategy: 'max_coverage', max_credits: maxCredits }, limit: maxCredits, }), }); if (!res.ok) throw new Error(`${res.status} ${await res.text()}`); const { id } = await res.json();
let page = await getSearch(id); while (!['completed', 'partial'].includes(page.status)) { if (page.status === 'failed') throw new Error(`search ${id} failed`); if (page.status === 'on_hold') console.warn('Out of credits. Top up and the search resumes.'); await sleep(15_000); page = await getSearch(id); } for (;;) { for (const job of page.data) await upsertJob(job); // the SQL upsert above if (!page.next_cursor) return; page = await getSearch(id, page.next_cursor); }}
await backfill(30, 5000);A finished search looks like this (illustrative, trimmed):
{ "id": "srch_2Vd9KqL4mN", "status": "completed", "jobs_found": 3184, "data": [{ "id": "job_01JC9F2K7NQ3XW5R8T1Y6M4H0C", "title": "Senior Data Engineer", "status": "open" }], "next_cursor": "cur_Lp0sR3", "metadata": { "status": "complete", "credits_charged": 3012, "jobs_already_paid": 172, "duplicates_merged": 1907 }}jobs_already_paid jobs were free: you had paid for them before. duplicates_merged provider records were folded into canonical jobs, also free.
Daily incremental
Section titled “Daily incremental”Once a day, search jobs first seen in the last 2 days and upsert them. The extra day is overlap: a late or failed run still catches everything. Overlap is free: jobs you already paid for come back at no charge and are counted in metadata.jobs_already_paid.
curl https://api.betterjobs.cc/v1/jobs/search \ -H "Authorization: Bearer $BETTERJOBS_API_KEY" \ -H "BetterJobs-Version: 2026-10-01" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: daily-2026-10-11-page-1" \ -d '{ "filters": { "title_or": ["Data Engineer", "Analytics Engineer"], "country_code_or": ["DE", "FR", "NL", "ES", "PL"], "posted_within_days": 2 }, "waterfall": { "strategy": "max_coverage", "max_credits": 100 }, "limit": 100 }'from datetime import date
def daily_sync(budget: int = 500) -> None: """Stops when the run has spent `budget` credits. max_credits caps each page only.""" cursor, page_no, spent = None, 1, 0 while spent < budget: body = { "filters": {**FILTERS, "posted_within_days": 2}, "waterfall": {"strategy": "max_coverage", "max_credits": min(100, budget - spent)}, "limit": 100, **({"cursor": cursor} if cursor else {}), } # Same key on a retry of the same page = never charged twice key = f"daily-{date.today().isoformat()}-page-{page_no}" resp = requests.post( f"{API}/jobs/search", headers={**HEADERS, "Idempotency-Key": key}, json=body, timeout=30 ) resp.raise_for_status() page = resp.json() for job in page["data"]: upsert_job(job) meta = page["metadata"] spent += meta["credits_charged"] print(meta["status"], meta["credits_charged"], "charged,", meta["jobs_already_paid"], "already paid") cursor, page_no = page["next_cursor"], page_no + 1 if cursor is None: return print(f"Run budget of {budget} credits reached. Resume tomorrow or raise the budget.")async function dailySync(budget = 500) { // Stops when the run has spent `budget` credits. max_credits caps each page only. let cursor: string | null = null; let pageNo = 1; let spent = 0; do { // Same key on a retry of the same page = never charged twice const key = `daily-${new Date().toISOString().slice(0, 10)}-page-${pageNo}`; const res = await fetch(`${API}/jobs/search`, { method: 'POST', headers: { ...headers, 'Idempotency-Key': key }, body: JSON.stringify({ filters: { ...filters, posted_within_days: 2 }, waterfall: { strategy: 'max_coverage', max_credits: Math.min(100, budget - spent) }, limit: 100, ...(cursor ? { cursor } : {}), }), }); if (!res.ok) throw new Error(`${res.status} ${await res.text()}`); const page = await res.json(); for (const job of page.data) await upsertJob(job); const { status, credits_charged, jobs_already_paid } = page.metadata; spent += credits_charged; console.log(status, credits_charged, 'charged,', jobs_already_paid, 'already paid'); cursor = page.next_cursor; pageNo += 1; } while (cursor && spent < budget); if (cursor) console.warn(`Run budget of ${budget} credits reached. Resume tomorrow or raise the budget.`);}If a daily window returns more jobs than you want to page through synchronously, run it as an async search instead, with the same filters and posted_within_days: 2.
Handle closed and updated jobs
Section titled “Handle closed and updated jobs”A daily search finds new jobs. It does not tell you when an old job closes. Create a search watch with the same filters and only the free lifecycle events:
curl https://api.betterjobs.cc/v1/watches \ -H "Authorization: Bearer $BETTERJOBS_API_KEY" \ -H "BetterJobs-Version: 2026-10-01" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: 6b0d8f2a-4c1e-4a73-9d5b-e7f9a1c3b508" \ -d '{ "type": "search", "filters": { "title_or": ["Data Engineer", "Analytics Engineer"], "country_code_or": ["DE", "FR", "NL", "ES", "PL"] }, "webhook_url": "https://hooks.northwind.example/betterjobs", "events": ["job.closed", "job.updated", "job.reposted"] }'def apply_event(event: dict) -> None: """Apply one event. Safe to call twice with the same event.""" job = event["data"].get("job") match event["type"]: case "job.updated" | "job.opened": upsert_job(job) # full Job case "job.closed": mark_closed(job["id"], job["closed_reason"]) # keep the row, set status and closed_reason case "job.reposted": set_repost_count(job["id"], job["repost_count"])
resp = requests.post( f"{API}/watches", headers={**HEADERS, "Idempotency-Key": str(uuid.uuid4())}, json={ "type": "search", "filters": FILTERS, "webhook_url": "https://hooks.northwind.example/betterjobs", "events": ["job.closed", "job.updated", "job.reposted"], }, timeout=30,)resp.raise_for_status()async function applyEvent(event: { type: string; data: { job?: any } }) { // Safe to call twice with the same event. const job = event.data.job; switch (event.type) { case 'job.updated': case 'job.opened': return upsertJob(job); // full Job case 'job.closed': return markClosed(job.id, job.closed_reason); // keep the row, set status and closed_reason case 'job.reposted': return setRepostCount(job.id, job.repost_count); }}
const res = await fetch(`${API}/watches`, { method: 'POST', headers: { ...headers, 'Idempotency-Key': crypto.randomUUID() }, body: JSON.stringify({ type: 'search', filters, webhook_url: 'https://hooks.northwind.example/betterjobs', events: ['job.closed', 'job.updated', 'job.reposted'], }),});if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);-- job.closed carries a partial job: update the columns, do not replace docUPDATE jobsSET status = 'closed', closed_reason = $2, doc = doc || jsonb_build_object('status', 'closed', 'closed_reason', $2)WHERE id = $1;Leaving job.opened out keeps the watch free: new jobs already arrive through the daily search. Webhook handling (signature check, event-id dedup) is on Detect hiring changes.
Recover from an outage
Section titled “Recover from an outage”Every event is also kept in GET /v1/events, oldest first. Store the last next_cursor you processed. After downtime, pass it as since and read until the feed is empty.
curl "https://api.betterjobs.cc/v1/events?since=cur_E5vB7n&limit=100" \ -H "Authorization: Bearer $BETTERJOBS_API_KEY" \ -H "BetterJobs-Version: 2026-10-01"def catch_up() -> None: since = load_state("events_cursor") # None on first run = oldest retained event while True: params = {"limit": 100, **({"since": since} if since else {})} resp = requests.get(f"{API}/events", headers=HEADERS, params=params, timeout=30) resp.raise_for_status() page = resp.json() for event in page["data"]: if not event_seen(event["id"]): # same id store as the webhook handler apply_event(event) mark_event_seen(event["id"]) if page["next_cursor"]: since = page["next_cursor"] save_state("events_cursor", since) # save after the page is applied if not page["data"] or page["next_cursor"] is None: returnasync function catchUp() { let since: string | null = await loadState('events_cursor'); // null = oldest retained event for (;;) { const qs = new URLSearchParams({ limit: '100', ...(since ? { since } : {}) }); const res = await fetch(`${API}/events?${qs}`, { headers }); if (!res.ok) throw new Error(`${res.status} ${await res.text()}`); const page = await res.json(); for (const event of page.data) { if (await eventSeen(event.id)) continue; // same id store as the webhook handler await applyEvent(event); await markEventSeen(event.id); } if (page.next_cursor) { since = page.next_cursor; await saveState('events_cursor', since); // save after the page is applied } if (page.data.length === 0 || !page.next_cursor) return; }}{ "data": [ { "id": "evt_3Fh8JkL2pQ", "type": "job.reposted", "created_at": "2026-10-11T07:30:00Z", "watch_id": "wat_6Np3QyR8tU", "data": { "job": { "id": "job_01JC2B7Y9MZQ4W8E1R6T3N5K0D", "status": "open", "repost_count": 2 } } } ], "next_cursor": "cur_E5vB7n"}Then run the daily incremental once. Jobs that opened during the outage come back; jobs you already have are free.
Three kinds of outage, three fixes:
| What was down | What you lost | Fix |
|---|---|---|
| Your webhook endpoint | Deliveries after the 24-hour retry window | GET /v1/events?since=<cursor>, or POST /v1/webhooks/replay per event |
| Your daily job | One or more daily runs | Re-run the daily search with posted_within_days covering the gap |
| An upstream provider | Jobs only that provider had | Responses say metadata.status: partial and list it in metadata.providers.failed. Re-run later; already-paid jobs are free |
Credit cost
Section titled “Credit cost”| Step | Cost |
|---|---|
dry_run estimate |
Free |
Backfill (POST /v1/searches) |
1 credit per unique job collected. Duplicates and already-paid jobs are free |
Reading search pages (GET /v1/searches/{id}) |
Free |
| Daily incremental | 1 credit per new unique job. The overlap day is free |
Watch with job.closed, job.updated, job.reposted |
Free |
GET /v1/events, replays |
Free |
GET /v1/jobs/{id} on a job you paid for |
Free |
Prove what you paid for with GET /v1/billing/ledger?job_id=job_...: re-reads show credits: 0 and reason: already_paid. See Credits and billing.
Pitfalls
Section titled “Pitfalls”posted_within_dayscounts fromfirst_seen_at. A job the employer posted weeks ago but a source found yesterday is in today’s window. That is what you want for sync.- Never key on a provider id.
sources[].provider_job_iddiffers per provider and a job can gain sources over time. Key on the canonicalid. - Do not delete closed jobs. Keep the row with
status: closedandclosed_reason. Searches skip closed jobs unless you sendinclude_closed: true. - Reuse the Idempotency-Key on retries. A retried
POSTwith the same key and body returns the first response and never charges twice. A new key per page per day, as above, is enough. See Idempotency. - Cap every scheduled run.
max_creditscaps one request, so for a paged sync run it caps each page only. Cap the run by summingmetadata.credits_charged, asdaily_syncdoes. An async search is one request, so itsmax_creditscaps the whole search.
- Detect hiring changes: webhook verification and event handling.
- Pagination, Async searches, Freshness and lifecycle.
- API reference: Create an async search, Get an async search, List events.