# Sync patterns

> How do I backfill, keep a daily copy in sync, and recover from outages?

Source: https://docs.betterjobs.cc/guides/sync-patterns/

Cost: 1 credit / unique jobJobs you already paid for come back free, so overlapping windows cost nothing extra.

**Goal:** a local table of every Data Engineer and Analytics Engineer job in five EU countries, filled once, kept current every day, and repaired after downtime.

Four patterns, one key: the canonical job `id`. It is stable across sources and requests, so every pattern ends in an upsert on `id`.

| Pattern                                              | Call                                                   | When                               |
| ---------------------------------------------------- | ------------------------------------------------------ | ---------------------------------- |
| [Initial backfill](#initial-backfill)                | `POST /v1/searches` (async)                            | Once, up to 10,000 jobs per search |
| [Daily incremental](#daily-incremental)              | `POST /v1/jobs/search` with `posted_within_days`       | Every day                          |
| [Lifecycle updates](#handle-closed-and-updated-jobs) | Watch with `job.closed`, `job.updated`, `job.reposted` | Continuous, free                   |
| [Outage recovery](#recover-from-an-outage)           | `GET /v1/events?since=`                                | After your side was down           |

## Your table

Store the full job as JSON plus the columns you query. Guard the upsert so older data never overwrites newer data.

```sql
CREATE TABLE jobs (
  id            text PRIMARY KEY,          -- canonical job id, job_...
  status        text NOT NULL,             -- open | closed
  closed_reason text,                      -- filled | expired | removed | unknown | NULL
  last_seen_at  timestamptz NOT NULL,
  doc           jsonb NOT NULL             -- the full Job object
);


CREATE TABLE sync_state (key text PRIMARY KEY, value text NOT NULL);


-- Upsert one job from a search page or a job.opened / job.updated event
INSERT INTO jobs (id, status, closed_reason, last_seen_at, doc)
VALUES ($1, $2, $3, $4, $5)
ON CONFLICT (id) DO UPDATE
SET status = EXCLUDED.status,
    closed_reason = EXCLUDED.closed_reason,
    last_seen_at = EXCLUDED.last_seen_at,
    doc = EXCLUDED.doc
WHERE jobs.last_seen_at <= EXCLUDED.last_seen_at;
```

## Initial backfill

1. **Estimate.** Send the filters to `POST /v1/jobs/search` with `dry_run: true`. Free. If `expected_unique_jobs_range.max` is above 10,000, split the backfill (see below).

2. **Create an async search.** `POST /v1/searches` with `limit` up to 10,000, `waterfall.max_credits` as a hard cap, and an `Idempotency-Key`. You get `202` and a `srch_...` id.

3. **Wait for it.** Pass `webhook_url` to receive `search.completed`, or poll `GET /v1/searches/{id}`. Branch on `status`, not on the HTTP code.

4. **Read every page.** Once `status` is `completed` or `partial`, `data` holds the first page. Pass `next_cursor` as `cursor` until it is `null`. Reading is free.

5. **Upsert each job on `id`.**

**curl**

```bash
# 1. Create the search (charged as jobs are collected)
curl https://api.betterjobs.cc/v1/searches \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: 0f6e2a94-7b3c-4d81-9e5a-c2b8d4f61a37" \
  -d '{
    "filters": {
      "title_or": ["Data Engineer", "Analytics Engineer"],
      "country_code_or": ["DE", "FR", "NL", "ES", "PL"],
      "posted_within_days": 30
    },
    "waterfall": { "strategy": "max_coverage", "max_credits": 5000 },
    "limit": 5000,
    "webhook_url": "https://hooks.northwind.example/betterjobs"
  }'


# 2. Check status / read pages (free). Add &cursor=<next_cursor> for the next page.
curl "https://api.betterjobs.cc/v1/searches/srch_2Vd9KqL4mN?limit=100" \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01"
```

**Python**

```python
import os
import time
import uuid


import requests


API = "https://api.betterjobs.cc/v1"
HEADERS = {
    "Authorization": f"Bearer {os.environ['BETTERJOBS_API_KEY']}",
    "BetterJobs-Version": "2026-10-01",
}
FILTERS = {
    "title_or": ["Data Engineer", "Analytics Engineer"],
    "country_code_or": ["DE", "FR", "NL", "ES", "PL"],
}




def backfill(days: int, max_credits: int) -> None:
    created = requests.post(
        f"{API}/searches",
        headers={**HEADERS, "Idempotency-Key": str(uuid.uuid4())},
        json={
            "filters": {**FILTERS, "posted_within_days": days},
            "waterfall": {"strategy": "max_coverage", "max_credits": max_credits},
            "limit": max_credits,
        },
        timeout=30,
    )
    created.raise_for_status()
    search_id = created.json()["id"]


    while True:
        page = get_search(search_id)
        if page["status"] in ("completed", "partial"):
            break
        if page["status"] == "failed":
            raise RuntimeError(f"search {search_id} failed")
        if page["status"] == "on_hold":
            print("Out of credits. Top up and the search resumes.")
        time.sleep(15)


    while True:
        for job in page["data"]:
            upsert_job(job)  # the SQL upsert above
        if page["next_cursor"] is None:
            return
        page = get_search(search_id, page["next_cursor"])




def get_search(search_id: str, cursor: str | None = None) -> dict:
    params = {"limit": 100, **({"cursor": cursor} if cursor else {})}
    resp = requests.get(f"{API}/searches/{search_id}", headers=HEADERS, params=params, timeout=30)
    resp.raise_for_status()
    return resp.json()




backfill(days=30, max_credits=5000)
```

**TypeScript**

```ts
const API = 'https://api.betterjobs.cc/v1';
const headers = {
  Authorization: `Bearer ${process.env.BETTERJOBS_API_KEY}`,
  'BetterJobs-Version': '2026-10-01',
  'Content-Type': 'application/json',
};
const filters = {
  title_or: ['Data Engineer', 'Analytics Engineer'],
  country_code_or: ['DE', 'FR', 'NL', 'ES', 'PL'],
};
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));


async function getSearch(id: string, cursor?: string) {
  const qs = new URLSearchParams({ limit: '100', ...(cursor ? { cursor } : {}) });
  const res = await fetch(`${API}/searches/${id}?${qs}`, { headers });
  if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
  return res.json();
}


async function backfill(days: number, maxCredits: number) {
  const res = await fetch(`${API}/searches`, {
    method: 'POST',
    headers: { ...headers, 'Idempotency-Key': crypto.randomUUID() },
    body: JSON.stringify({
      filters: { ...filters, posted_within_days: days },
      waterfall: { strategy: 'max_coverage', max_credits: maxCredits },
      limit: maxCredits,
    }),
  });
  if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
  const { id } = await res.json();


  let page = await getSearch(id);
  while (!['completed', 'partial'].includes(page.status)) {
    if (page.status === 'failed') throw new Error(`search ${id} failed`);
    if (page.status === 'on_hold') console.warn('Out of credits. Top up and the search resumes.');
    await sleep(15_000);
    page = await getSearch(id);
  }
  for (;;) {
    for (const job of page.data) await upsertJob(job); // the SQL upsert above
    if (!page.next_cursor) return;
    page = await getSearch(id, page.next_cursor);
  }
}


await backfill(30, 5000);
```

A finished search looks like this (illustrative, trimmed):

```json
{
  "id": "srch_2Vd9KqL4mN",
  "status": "completed",
  "jobs_found": 3184,
  "data": [{ "id": "job_01JC9F2K7NQ3XW5R8T1Y6M4H0C", "title": "Senior Data Engineer", "status": "open" }],
  "next_cursor": "cur_Lp0sR3",
  "metadata": {
    "status": "complete",
    "credits_charged": 3012,
    "jobs_already_paid": 172,
    "duplicates_merged": 1907
  }
}
```

`jobs_already_paid` jobs were free: you had paid for them before. `duplicates_merged` provider records were folded into canonical jobs, also free.

> More than 10,000 jobs
>
> One async search returns up to 10,000 jobs. Split larger backfills by a filter that partitions the set, for example one search per `country_code_or` value or per `job_family_or` value. Overlap between slices is safe: a job you already paid for comes back free, and the upsert on `id` keeps one row.

## Daily incremental

Once a day, search jobs first seen in the last **2** days and upsert them. The extra day is overlap: a late or failed run still catches everything. Overlap is free: jobs you already paid for come back at no charge and are counted in `metadata.jobs_already_paid`.

**curl**

```bash
curl https://api.betterjobs.cc/v1/jobs/search \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: daily-2026-10-11-page-1" \
  -d '{
    "filters": {
      "title_or": ["Data Engineer", "Analytics Engineer"],
      "country_code_or": ["DE", "FR", "NL", "ES", "PL"],
      "posted_within_days": 2
    },
    "waterfall": { "strategy": "max_coverage", "max_credits": 100 },
    "limit": 100
  }'
```

**Python**

```python
from datetime import date




def daily_sync(budget: int = 500) -> None:
    """Stops when the run has spent `budget` credits. max_credits caps each page only."""
    cursor, page_no, spent = None, 1, 0
    while spent < budget:
        body = {
            "filters": {**FILTERS, "posted_within_days": 2},
            "waterfall": {"strategy": "max_coverage", "max_credits": min(100, budget - spent)},
            "limit": 100,
            **({"cursor": cursor} if cursor else {}),
        }
        # Same key on a retry of the same page = never charged twice
        key = f"daily-{date.today().isoformat()}-page-{page_no}"
        resp = requests.post(
            f"{API}/jobs/search", headers={**HEADERS, "Idempotency-Key": key}, json=body, timeout=30
        )
        resp.raise_for_status()
        page = resp.json()
        for job in page["data"]:
            upsert_job(job)
        meta = page["metadata"]
        spent += meta["credits_charged"]
        print(meta["status"], meta["credits_charged"], "charged,", meta["jobs_already_paid"], "already paid")
        cursor, page_no = page["next_cursor"], page_no + 1
        if cursor is None:
            return
    print(f"Run budget of {budget} credits reached. Resume tomorrow or raise the budget.")
```

**TypeScript**

```ts
async function dailySync(budget = 500) {
  // Stops when the run has spent `budget` credits. max_credits caps each page only.
  let cursor: string | null = null;
  let pageNo = 1;
  let spent = 0;
  do {
    // Same key on a retry of the same page = never charged twice
    const key = `daily-${new Date().toISOString().slice(0, 10)}-page-${pageNo}`;
    const res = await fetch(`${API}/jobs/search`, {
      method: 'POST',
      headers: { ...headers, 'Idempotency-Key': key },
      body: JSON.stringify({
        filters: { ...filters, posted_within_days: 2 },
        waterfall: { strategy: 'max_coverage', max_credits: Math.min(100, budget - spent) },
        limit: 100,
        ...(cursor ? { cursor } : {}),
      }),
    });
    if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
    const page = await res.json();
    for (const job of page.data) await upsertJob(job);
    const { status, credits_charged, jobs_already_paid } = page.metadata;
    spent += credits_charged;
    console.log(status, credits_charged, 'charged,', jobs_already_paid, 'already paid');
    cursor = page.next_cursor;
    pageNo += 1;
  } while (cursor && spent < budget);
  if (cursor) console.warn(`Run budget of ${budget} credits reached. Resume tomorrow or raise the budget.`);
}
```

> Use the same strategy as the backfill
>
> `cheapest_first` stops once a page is full, so it is the wrong choice for a complete copy. Use `max_coverage` for backfill and daily runs alike, so both see the same sources. See [Choose a strategy](https://docs.betterjobs.cc/guides/choose-a-strategy.md).

If a daily window returns more jobs than you want to page through synchronously, run it as an async search instead, with the same filters and `posted_within_days: 2`.

## Handle closed and updated jobs

A daily search finds new jobs. It does not tell you when an old job closes. Create a search watch with the same filters and only the free lifecycle events:

**curl**

```bash
curl https://api.betterjobs.cc/v1/watches \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: 6b0d8f2a-4c1e-4a73-9d5b-e7f9a1c3b508" \
  -d '{
    "type": "search",
    "filters": {
      "title_or": ["Data Engineer", "Analytics Engineer"],
      "country_code_or": ["DE", "FR", "NL", "ES", "PL"]
    },
    "webhook_url": "https://hooks.northwind.example/betterjobs",
    "events": ["job.closed", "job.updated", "job.reposted"]
  }'
```

**Python**

```python
def apply_event(event: dict) -> None:
    """Apply one event. Safe to call twice with the same event."""
    job = event["data"].get("job")
    match event["type"]:
        case "job.updated" | "job.opened":
            upsert_job(job)  # full Job
        case "job.closed":
            mark_closed(job["id"], job["closed_reason"])  # keep the row, set status and closed_reason
        case "job.reposted":
            set_repost_count(job["id"], job["repost_count"])




resp = requests.post(
    f"{API}/watches",
    headers={**HEADERS, "Idempotency-Key": str(uuid.uuid4())},
    json={
        "type": "search",
        "filters": FILTERS,
        "webhook_url": "https://hooks.northwind.example/betterjobs",
        "events": ["job.closed", "job.updated", "job.reposted"],
    },
    timeout=30,
)
resp.raise_for_status()
```

**TypeScript**

```ts
async function applyEvent(event: { type: string; data: { job?: any } }) {
  // Safe to call twice with the same event.
  const job = event.data.job;
  switch (event.type) {
    case 'job.updated':
    case 'job.opened':
      return upsertJob(job); // full Job
    case 'job.closed':
      return markClosed(job.id, job.closed_reason); // keep the row, set status and closed_reason
    case 'job.reposted':
      return setRepostCount(job.id, job.repost_count);
  }
}


const res = await fetch(`${API}/watches`, {
  method: 'POST',
  headers: { ...headers, 'Idempotency-Key': crypto.randomUUID() },
  body: JSON.stringify({
    type: 'search',
    filters,
    webhook_url: 'https://hooks.northwind.example/betterjobs',
    events: ['job.closed', 'job.updated', 'job.reposted'],
  }),
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
```

```sql
-- job.closed carries a partial job: update the columns, do not replace doc
UPDATE jobs
SET status = 'closed',
    closed_reason = $2,
    doc = doc || jsonb_build_object('status', 'closed', 'closed_reason', $2)
WHERE id = $1;
```

Leaving `job.opened` out keeps the watch free: new jobs already arrive through the daily search. Webhook handling (signature check, event-id dedup) is on [Detect hiring changes](https://docs.betterjobs.cc/guides/detect-hiring-changes.md#handle-deliveries).

> No watch? Re-check stale jobs
>
> Without a watch, re-check open jobs whose `last_seen_at` is old with `GET /v1/jobs/{id}`. It is free for jobs you already paid for and returns the current `status` and `closed_reason`.

## Recover from an outage

Every event is also kept in `GET /v1/events`, oldest first. Store the last `next_cursor` you processed. After downtime, pass it as `since` and read until the feed is empty.

**curl**

```bash
curl "https://api.betterjobs.cc/v1/events?since=cur_E5vB7n&limit=100" \
  -H "Authorization: Bearer $BETTERJOBS_API_KEY" \
  -H "BetterJobs-Version: 2026-10-01"
```

**Python**

```python
def catch_up() -> None:
    since = load_state("events_cursor")  # None on first run = oldest retained event
    while True:
        params = {"limit": 100, **({"since": since} if since else {})}
        resp = requests.get(f"{API}/events", headers=HEADERS, params=params, timeout=30)
        resp.raise_for_status()
        page = resp.json()
        for event in page["data"]:
            if not event_seen(event["id"]):  # same id store as the webhook handler
                apply_event(event)
                mark_event_seen(event["id"])
        if page["next_cursor"]:
            since = page["next_cursor"]
            save_state("events_cursor", since)  # save after the page is applied
        if not page["data"] or page["next_cursor"] is None:
            return
```

**TypeScript**

```ts
async function catchUp() {
  let since: string | null = await loadState('events_cursor'); // null = oldest retained event
  for (;;) {
    const qs = new URLSearchParams({ limit: '100', ...(since ? { since } : {}) });
    const res = await fetch(`${API}/events?${qs}`, { headers });
    if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
    const page = await res.json();
    for (const event of page.data) {
      if (await eventSeen(event.id)) continue; // same id store as the webhook handler
      await applyEvent(event);
      await markEventSeen(event.id);
    }
    if (page.next_cursor) {
      since = page.next_cursor;
      await saveState('events_cursor', since); // save after the page is applied
    }
    if (page.data.length === 0 || !page.next_cursor) return;
  }
}
```

```json
{
  "data": [
    { "id": "evt_3Fh8JkL2pQ", "type": "job.reposted", "created_at": "2026-10-11T07:30:00Z", "watch_id": "wat_6Np3QyR8tU", "data": { "job": { "id": "job_01JC2B7Y9MZQ4W8E1R6T3N5K0D", "status": "open", "repost_count": 2 } } }
  ],
  "next_cursor": "cur_E5vB7n"
}
```

Then run the daily incremental once. Jobs that opened during the outage come back; jobs you already have are free.

Three kinds of outage, three fixes:

| What was down         | What you lost                             | Fix                                                                                                                           |
| --------------------- | ----------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| Your webhook endpoint | Deliveries after the 24-hour retry window | `GET /v1/events?since=<cursor>`, or `POST /v1/webhooks/replay` per event                                                      |
| Your daily job        | One or more daily runs                    | Re-run the daily search with `posted_within_days` covering the gap                                                            |
| An upstream provider  | Jobs only that provider had               | Responses say `metadata.status: partial` and list it in `metadata.providers.failed`. Re-run later; already-paid jobs are free |

## Credit cost

| Step                                                   | Cost                                                                         |
| ------------------------------------------------------ | ---------------------------------------------------------------------------- |
| `dry_run` estimate                                     | Free                                                                         |
| Backfill (`POST /v1/searches`)                         | 1 credit per unique job collected. Duplicates and already-paid jobs are free |
| Reading search pages (`GET /v1/searches/{id}`)         | Free                                                                         |
| Daily incremental                                      | 1 credit per new unique job. The overlap day is free                         |
| Watch with `job.closed`, `job.updated`, `job.reposted` | Free                                                                         |
| `GET /v1/events`, replays                              | Free                                                                         |
| `GET /v1/jobs/{id}` on a job you paid for              | Free                                                                         |

Prove what you paid for with `GET /v1/billing/ledger?job_id=job_...`: re-reads show `credits: 0` and `reason: already_paid`. See [Credits and billing](https://docs.betterjobs.cc/concepts/credits-and-billing.md).

## Pitfalls

> Branch on status, not on HTTP 200
>
> `GET /v1/searches/{id}` returns `200` while the search is `queued`, `running` or `on_hold`. Only `completed` and `partial` carry results. `on_hold` means out of credits; the search resumes after a top-up. See [Async searches](https://docs.betterjobs.cc/platform/async-searches.md).

- **`posted_within_days` counts from `first_seen_at`.** A job the employer posted weeks ago but a source found yesterday is in today’s window. That is what you want for sync.
- **Never key on a provider id.** `sources[].provider_job_id` differs per provider and a job can gain sources over time. Key on the canonical `id`.
- **Do not delete closed jobs.** Keep the row with `status: closed` and `closed_reason`. Searches skip closed jobs unless you send `include_closed: true`.
- **Reuse the Idempotency-Key on retries.** A retried `POST` with the same key and body returns the first response and never charges twice. A new key per page per day, as above, is enough. See [Idempotency](https://docs.betterjobs.cc/platform/idempotency.md).
- **Cap every scheduled run.** `max_credits` caps one request, so for a paged sync run it caps each page only. Cap the run by summing `metadata.credits_charged`, as `daily_sync` does. An async search is one request, so its `max_credits` caps the whole search.

## Next

- [Detect hiring changes](https://docs.betterjobs.cc/guides/detect-hiring-changes.md): webhook verification and event handling.
- [Pagination](https://docs.betterjobs.cc/platform/pagination.md), [Async searches](https://docs.betterjobs.cc/platform/async-searches.md), [Freshness and lifecycle](https://docs.betterjobs.cc/concepts/freshness-and-lifecycle.md).
- API reference: [Create an async search](/api/operations/createsearch/), [Get an async search](/api/operations/getsearch/), [List events](/api/operations/listevents/).
