Gigflow docs
Engineering

ADR-0015: Vacancy data model

Store each model classification once and derive typed vacancy columns from it.

Status: Proposed. Implementation available for review.

Date: 2026-09-28

Context

The classify stage returns quotes and structured records (#253), but the vacancy tables still had the columns of the old flat extraction. The upsert stored null in seniority, employment type, work model, salary, and skills. It kept the model output only in vacancies.content, which the next change overwrote, and it did not keep the output of rejected pages. It wrote about 30 vacancy_statements rows on each revision, and the page Markdown twice: in vacancy_source_versions.source_markdown and in vacancies.description_raw.

The first full harvest classifies every vacancy page with three model calls. Model output that is lost must be paid for again. Data that is derived from it must not need a new model call.

Decision

Each kind of data has one table:

DataTableHolds
Pagevacancy_source_versionsThe object storage key of the harvested document, its hash, the page type, and the classification. The key contains the SHA-256 of the document, so an unchanged page is stored once over all harvests.
Model outputvacancy_classificationsThe output, the SHA-256 of the prepared Markdown, the classifier version, the model, and the token totals
FactsvacanciesTyped columns derived from the current classification
Evidencevacancy_statementsThe source quotes of the current vacancy, with type, subtype, level, and classification path
Placesvacancy_locationsGeocoded locations; is_primary marks the first
Vectorsvacancy_embeddingsNo change

The vacancy is the central record. The API, Admin, and matching read vacancies and its statements, locations, and embedding. Only the pipeline reads vacancy_classifications: to reuse output, to track cost, and to rebuild derived data.

The classify job prepares the Markdown and looks up a classification with the same input hash and classifier version. When it finds one, it uses it without a model call. Otherwise it calls the model and stores the result, also for a page without a type, before the job fails: the next harvest of that page makes no model call. The job returns classificationId, the creation time of the classification as classifiedAt, and the output, not the Markdown. The Markdown is not stored: it can be prepared again from the HTML in object storage. classifyConfig.version increases when the prompts, the schema, or the model change.

The upsert derives the vacancy columns from the classification with these rules:

  • One-to-one fields are columns: title, job_function, role_category, seniority_level, references, and company_name_raw from the hiring organization. language is the ISO 639-1 code of the vacancy text, or null for another value.
  • The model can miss a title that the page states. When the classification of an open job page has no title, title takes the longest part of the page title, without the parts that name an organization of the classification or a label of the host name: "Bizzy - Internship" gives "Internship". A dash separates parts only between spaces. The stored classification keeps the model output, and the upsert logs each such title, so the rate is known. A generic page title, such as "Careers | Acme", gives a generic title.
  • posted_date and application_deadline are the dates that the page states. A date without a year gets the year that puts it nearest to classifiedAt, and a publication date is not later than it. A month without a day is the first day for a publication date and the last day for a deadline. A date without a month, or one that does not exist, is null. A rebuild from the stored classification gives the same dates. start_date follows the same rule, with the first day for a month without a day. is_immediate_start is true for a start as soon as possible and null when no start is stated.
  • residence_country_codes holds the countries that eligibility and remote-work conditions name. A region, such as the EU, has no code.
  • An array holds the distinct known values of all options: engagement_arrangements, contract_terms, and work_arrangement_modes. Null means that no value is known. No enum has an unknown value.
  • Workloads are alternatives. minimum_hours_per_week and maximum_hours_per_week cover all options, and so do minimum_days_per_week and maximum_days_per_week. A bound is null when an option does not state it in that unit.
  • Work arrangements are statements about one vacancy. maximum_remote_days_per_week is the highest stated value.
  • Pay comes from the first salary or contract rate: minimum_pay, maximum_pay, pay_currency, and pay_period.
  • The hard filters of ADR-0014 come from requirements with level required and no alternative. required_languages holds the ISO 639-1 code of each such language requirement that names one language. required_education_level and required_education_fields hold the ISCED level and ISCED-F fields of the highest such degree, because a higher level meets a lower requirement. CEFR levels and the other requirements stay in the classification and the statements.
  • Each quote of the classification becomes a statement, in a fixed order: requirements, responsibilities, work conditions, work arrangements, engagements and their workloads, compensation, employment terms, benefits, and context. Contacts, organizations, and locations are not statements.

The database generates the search vectors. Each row indexes its own text: vacancies.fts_vector the title (A) and the job function and company name (B), and vacancy_statements.fts_vector the quote, with weight C for requirements and responsibilities and D for the other types. A vacancy matches a search on its own row or on any of its statements, and a statement match shows which quote contains the term.

A vacancy with the same classification is unchanged, also when the page markup changed. Then only last_seen_at changes, and the statement IDs stay the same. A new classification replaces the statements.

Harvests and expiry

A harvest keeps one row over the retries of its job. The first attempt of a configuration harvest job stores the harvest ID in the job data, and a retry continues that row. One configuration runs at most one harvest: a harvest start locks the configuration row, and a partial unique index allows one running harvest per configuration. A running harvest that started more than six hours ago stopped without completion, so the next start marks it failed. Each source gets one ingestion per harvest, with the job ID vacancy-<source ID>-<harvest ID>. When a retry finds other content for a source that it already scheduled, the first ingestion stays.

A source expires when:

  • Two consecutive complete harvests of its configuration do not list it. missed_harvest_count counts these harvests, and a listing resets it. An incomplete harvest expires nothing. A harvest is incomplete when it stops at a safety limit, or when it finds less than half of the previous item count of 20 or more.
  • Its page is a closed job, an error page, a job listing, or not a job page. A page that requires a login does not expire it.

A sitemap that cannot be read hides its sources, so it must not look like a sitemap without jobs. When no sitemap gives URLs and one could not be read, or the discovery timed out, the harvest fails and retries. When a sitemap that a sitemap index or robots.txt lists cannot be read, the discovery times out, or the limit of 25 sitemaps skips a listed one, the harvest is incomplete. A failed guess, such as a missing /sitemap.xml, matters only when no sitemap gives URLs.

A vacancy expires when all its sources have expired, and a closed job page closes it. When a harvest lists an expired source again, its ingestion makes the vacancy active.

Each upsert adds one to a count of its harvest: new_count, changed_count, or unchanged_count for a job page, and rejected_count for another page. The last failed attempt of a flow adds one to failed_count, and writes the stage, the error message, and the time on the source (last_error_stage, last_error_message, and last_error_at). The message of a pipeline error starts with its code, such as [EXTRACTION_FAILED]. A later stored page clears the error. The counts grow while the ingestions run, also after the harvest row is complete. A retry of an upsert after its commit can count a source twice. A job that fails because it stalled too often is not counted.

The classify, geocode, embed, and upsert jobs retry by error category: a provider rate limit after 5 seconds, up to 5 times; a server error or timeout after 2 seconds, up to 3 times; an unavailable geocoding service after 30 seconds, up to 3 times. A request that the provider rejects with another 4xx status is not repeated. The harvest job keeps the fast default retries, because it holds the admission of its host while it waits. A site that answers 429 blocks its host for the wait of its Retry-After header, at least 30 seconds and at most one day, and for 5 minutes without the header.

Listing and sitemap URLs get one form: without fragment, credentials, trailing slash, and utm_* parameters, and with sorted query parameters. A page therefore has one source key in all strategies.

The pods of the vacancy workers give active jobs a termination grace period: 300 seconds for classify, 150 for embed, and 60 for the other stages. The shutdown timeout of each worker is 20 seconds shorter. A job that a deploy stops runs again, and the third stop fails it.

The database enums take their values from the tuples in @gigflow/engine-jobs/vacancy/schema, and the Engine API builds its enums from the database enums. The API list filters by seniority level, engagement arrangement, contract term, work arrangement mode, the country of any location, and a range of publication dates. The detail returns the statements.

Alternatives

AlternativeReason for this decision
A table per classification array, such as engagements and compensationsFive more tables, a delete and insert per table on each change, and a join per filter. The filter columns on vacancies are faster for nearest-neighbor search with filters.
Only the JSON output, with JSONPath queriesWeak typing and queries that are hard to read.
Statements per revision (the earlier shape)Every changed page hash added about 30 rows, also for the same content.
Statements on the classificationIDs that never change, but the vacancy is no longer the central record, and unused classifications keep their statements. An assessment of changed text is stale anyway.
A search vector that the upsert writesThe generated columns stay consistent without application code.
Store the MarkdownIt is large, and the HTML in object storage gives it again.

Exact checks of paired options, such as a freelance day rate for 5 days or an employment for 4 days, read the classification output in code. The SQL filters exclude a vacancy only on a conflict with the derived columns.

Consequences

Migration 0025 deletes all vacancies and source versions: they come from test harvests and have no classification. Migration 0026 adds the seniority level, the language, and the two dates, and replaces posted_at and closes_at, which were always null. Migration 0027 adds the required languages and the required degree, and migration 0028 the days per week, the start, and the residence countries. Sources, configurations, and harvests stay, so the next harvest classifies every source again. Pause vacancy harvests during the deploy; flows of the old pipeline fail.

Migration 0029 adds the rejected count, the missed-harvest count, and the index of running harvests. Before it builds the index, it marks the older running harvests of a configuration failed.

Migration 0030 adds the error columns to vacancy_sources. Deploy the vacancy workers with the flow producers: a flow job with the classified backoff fails on a worker without that strategy. Sources whose URL form changes get a new source once, and the old source expires after two complete harvests.

An error page expires its source at once. A temporary error page therefore expires a vacancy until the next harvest lists the source and its ingestion accepts the page again.

A new derived column needs only code and a rebuild from vacancy_classifications, without a model call.

Match assessments (ADR-0003) refer to statements with a cascading foreign key. When the classification of a vacancy changes, its statements and their assessments are replaced, and matching assesses the vacancy again.

Verification

Set TEST_DATABASE_URL to a migrated local database whose name contains test. From the repository root, with Redis, PostgreSQL, and MinIO from dev/cli/dev up engine, run each stage suite:

bun --env-file=apps/engine/apps/workers/vacancy/classify/.env \
  test --timeout 30000 apps/engine/apps/workers/vacancy/classify/e2e
bun --env-file=apps/engine/apps/workers/vacancy/upsert/.env \
  test --timeout 30000 apps/engine/apps/workers/vacancy/upsert/e2e

The classify test replaces only the model call. The upsert test runs the real worker and database. Reports are written to .context/verification/vacancy-classify/ and .context/verification/vacancy-upsert/.

The configuration harvest test runs the real worker, database, queues, and storage, and replaces only the browser harvester. It checks the harvest retry, the running-harvest lock, the expiry, and the item-drop guard. Its report is written to .context/verification/configuration-harvest/:

bun --env-file=apps/engine/apps/workers/configuration/harvest/.env \
  test --timeout 30000 apps/engine/apps/workers/configuration/harvest/e2e

The ingestion flow test runs the five vacancy workers for every HTML eval fixture, from admission to upsert. It uses the local browser service, the real model and embedding provider, and a Pelias fixture, because local development has no geocoding service. It makes paid model calls, so it runs only when VACANCY_FLOW_E2E is true:

VACANCY_FLOW_E2E=true bun --env-file=apps/engine/apps/workers/vacancy/upsert/.env \
  test --timeout 3600000 \
  apps/engine/apps/workers/vacancy/upsert/e2e/ingestion-flow.e2e.test.ts

Its report, with the outcome, facts, and token usage of each page, is written to .context/verification/vacancy-flow/.

On this page