Skip to content
API v1 preview — endpoints and fields may change before general availability.

Sync patterns

How do I backfill, keep a daily copy in sync, and recover from outages?

View .md
Cost: 1 credit / unique jobJobs you already paid for come back free, so overlapping windows cost nothing extra.

Goal: a local table of every Data Engineer and Analytics Engineer job in five EU countries, filled once, kept current every day, and repaired after downtime.

Four patterns, one key: the canonical job id. It is stable across sources and requests, so every pattern ends in an upsert on id.

Pattern Call When
Initial backfill POST /v1/searches (async) Once, up to 10,000 jobs per search
Daily incremental POST /v1/jobs/search with posted_within_days Every day
Lifecycle updates Watch with job.closed, job.updated, job.reposted Continuous, free
Outage recovery GET /v1/events?since= After your side was down

Store the full job as JSON plus the columns you query. Guard the upsert so older data never overwrites newer data.

CREATE TABLE jobs (
id text PRIMARY KEY, -- canonical job id, job_...
status text NOT NULL, -- open | closed
closed_reason text, -- filled | expired | removed | unknown | NULL
last_seen_at timestamptz NOT NULL,
doc jsonb NOT NULL -- the full Job object
);
CREATE TABLE sync_state (key text PRIMARY KEY, value text NOT NULL);
-- Upsert one job from a search page or a job.opened / job.updated event
INSERT INTO jobs (id, status, closed_reason, last_seen_at, doc)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (id) DO UPDATE
SET status = EXCLUDED.status,
closed_reason = EXCLUDED.closed_reason,
last_seen_at = EXCLUDED.last_seen_at,
doc = EXCLUDED.doc
WHERE jobs.last_seen_at <= EXCLUDED.last_seen_at;
  1. Estimate. Send the filters to POST /v1/jobs/search with dry_run: true. Free. If expected_unique_jobs_range.max is above 10,000, split the backfill (see below).

  2. Create an async search. POST /v1/searches with limit up to 10,000, waterfall.max_credits as a hard cap, and an Idempotency-Key. You get 202 and a srch_... id.

  3. Wait for it. Pass webhook_url to receive search.completed, or poll GET /v1/searches/{id}. Branch on status, not on the HTTP code.

  4. Read every page. Once status is completed or partial, data holds the first page. Pass next_cursor as cursor until it is null. Reading is free.

  5. Upsert each job on id.

Terminal window
# 1. Create the search (charged as jobs are collected)
curl https://api.betterjobs.cc/v1/searches \
-H "Authorization: Bearer $BETTERJOBS_API_KEY" \
-H "BetterJobs-Version: 2026-10-01" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 0f6e2a94-7b3c-4d81-9e5a-c2b8d4f61a37" \
-d '{
"filters": {
"title_or": ["Data Engineer", "Analytics Engineer"],
"country_code_or": ["DE", "FR", "NL", "ES", "PL"],
"posted_within_days": 30
},
"waterfall": { "strategy": "max_coverage", "max_credits": 5000 },
"limit": 5000,
"webhook_url": "https://hooks.northwind.example/betterjobs"
}'
# 2. Check status / read pages (free). Add &cursor=<next_cursor> for the next page.
curl "https://api.betterjobs.cc/v1/searches/srch_2Vd9KqL4mN?limit=100" \
-H "Authorization: Bearer $BETTERJOBS_API_KEY" \
-H "BetterJobs-Version: 2026-10-01"

A finished search looks like this (illustrative, trimmed):

{
"id": "srch_2Vd9KqL4mN",
"status": "completed",
"jobs_found": 3184,
"data": [{ "id": "job_01JC9F2K7NQ3XW5R8T1Y6M4H0C", "title": "Senior Data Engineer", "status": "open" }],
"next_cursor": "cur_Lp0sR3",
"metadata": {
"status": "complete",
"credits_charged": 3012,
"jobs_already_paid": 172,
"duplicates_merged": 1907
}
}

jobs_already_paid jobs were free: you had paid for them before. duplicates_merged provider records were folded into canonical jobs, also free.

Once a day, search jobs first seen in the last 2 days and upsert them. The extra day is overlap: a late or failed run still catches everything. Overlap is free: jobs you already paid for come back at no charge and are counted in metadata.jobs_already_paid.

Terminal window
curl https://api.betterjobs.cc/v1/jobs/search \
-H "Authorization: Bearer $BETTERJOBS_API_KEY" \
-H "BetterJobs-Version: 2026-10-01" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: daily-2026-10-11-page-1" \
-d '{
"filters": {
"title_or": ["Data Engineer", "Analytics Engineer"],
"country_code_or": ["DE", "FR", "NL", "ES", "PL"],
"posted_within_days": 2
},
"waterfall": { "strategy": "max_coverage", "max_credits": 100 },
"limit": 100
}'

If a daily window returns more jobs than you want to page through synchronously, run it as an async search instead, with the same filters and posted_within_days: 2.

A daily search finds new jobs. It does not tell you when an old job closes. Create a search watch with the same filters and only the free lifecycle events:

Terminal window
curl https://api.betterjobs.cc/v1/watches \
-H "Authorization: Bearer $BETTERJOBS_API_KEY" \
-H "BetterJobs-Version: 2026-10-01" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 6b0d8f2a-4c1e-4a73-9d5b-e7f9a1c3b508" \
-d '{
"type": "search",
"filters": {
"title_or": ["Data Engineer", "Analytics Engineer"],
"country_code_or": ["DE", "FR", "NL", "ES", "PL"]
},
"webhook_url": "https://hooks.northwind.example/betterjobs",
"events": ["job.closed", "job.updated", "job.reposted"]
}'
-- job.closed carries a partial job: update the columns, do not replace doc
UPDATE jobs
SET status = 'closed',
closed_reason = $2,
doc = doc || jsonb_build_object('status', 'closed', 'closed_reason', $2)
WHERE id = $1;

Leaving job.opened out keeps the watch free: new jobs already arrive through the daily search. Webhook handling (signature check, event-id dedup) is on Detect hiring changes.

Every event is also kept in GET /v1/events, oldest first. Store the last next_cursor you processed. After downtime, pass it as since and read until the feed is empty.

Terminal window
curl "https://api.betterjobs.cc/v1/events?since=cur_E5vB7n&limit=100" \
-H "Authorization: Bearer $BETTERJOBS_API_KEY" \
-H "BetterJobs-Version: 2026-10-01"
{
"data": [
{ "id": "evt_3Fh8JkL2pQ", "type": "job.reposted", "created_at": "2026-10-11T07:30:00Z", "watch_id": "wat_6Np3QyR8tU", "data": { "job": { "id": "job_01JC2B7Y9MZQ4W8E1R6T3N5K0D", "status": "open", "repost_count": 2 } } }
],
"next_cursor": "cur_E5vB7n"
}

Then run the daily incremental once. Jobs that opened during the outage come back; jobs you already have are free.

Three kinds of outage, three fixes:

What was down What you lost Fix
Your webhook endpoint Deliveries after the 24-hour retry window GET /v1/events?since=<cursor>, or POST /v1/webhooks/replay per event
Your daily job One or more daily runs Re-run the daily search with posted_within_days covering the gap
An upstream provider Jobs only that provider had Responses say metadata.status: partial and list it in metadata.providers.failed. Re-run later; already-paid jobs are free
Step Cost
dry_run estimate Free
Backfill (POST /v1/searches) 1 credit per unique job collected. Duplicates and already-paid jobs are free
Reading search pages (GET /v1/searches/{id}) Free
Daily incremental 1 credit per new unique job. The overlap day is free
Watch with job.closed, job.updated, job.reposted Free
GET /v1/events, replays Free
GET /v1/jobs/{id} on a job you paid for Free

Prove what you paid for with GET /v1/billing/ledger?job_id=job_...: re-reads show credits: 0 and reason: already_paid. See Credits and billing.

  • posted_within_days counts from first_seen_at. A job the employer posted weeks ago but a source found yesterday is in today’s window. That is what you want for sync.
  • Never key on a provider id. sources[].provider_job_id differs per provider and a job can gain sources over time. Key on the canonical id.
  • Do not delete closed jobs. Keep the row with status: closed and closed_reason. Searches skip closed jobs unless you send include_closed: true.
  • Reuse the Idempotency-Key on retries. A retried POST with the same key and body returns the first response and never charges twice. A new key per page per day, as above, is enough. See Idempotency.
  • Cap every scheduled run. max_credits caps one request, so for a paged sync run it caps each page only. Cap the run by summing metadata.credits_charged, as daily_sync does. An async search is one request, so its max_credits caps the whole search.