ADR-0015: Vacancy data model
Store each model classification once and derive typed vacancy columns from it.
Status: Proposed. Implementation available for review.
Date: 2026-09-28
Context
The classify stage returns quotes and structured records (#253), but the
vacancy tables still had the columns of the old flat extraction. The upsert
stored null in seniority, employment type, work model, salary, and skills. It
kept the model output only in vacancies.content, which the next change
overwrote, and it did not keep the output of rejected pages. It wrote about 30
vacancy_statements rows on each revision, and the page Markdown twice: in
vacancy_source_versions.source_markdown and in vacancies.description_raw.
The first full harvest classifies every vacancy page with three model calls. Model output that is lost must be paid for again. Data that is derived from it must not need a new model call.
Decision
Each kind of data has one table:
| Data | Table | Holds |
|---|---|---|
| Page | vacancy_source_versions | The object storage key of the harvested document, its hash, the page type, and the classification. The key contains the SHA-256 of the document, so an unchanged page is stored once over all harvests. |
| Model output | vacancy_classifications | The output, the SHA-256 of the prepared Markdown, the classifier version, the model, and the token totals |
| Facts | vacancies | Typed columns derived from the current classification |
| Evidence | vacancy_statements | The source quotes of the current vacancy, with type, subtype, level, and classification path |
| Places | vacancy_locations | Geocoded locations; is_primary marks the first |
| Vectors | vacancy_embeddings | No change |
The vacancy is the central record. The API, Admin, and matching read
vacancies and its statements, locations, and embedding. Only the pipeline
reads vacancy_classifications: to reuse output, to track cost, and to rebuild
derived data.
The classify job prepares the Markdown and looks up a classification with the
same input hash and classifier version. When it finds one, it uses it without a
model call. Otherwise it calls the model and stores the result, also for a page
without a type, before the job fails: the next harvest of that page makes no
model call. The job returns
classificationId, the creation time of the classification as classifiedAt,
and the output, not the Markdown. The Markdown is not stored:
it can be prepared again from the HTML in object storage. classifyConfig.version
increases when the prompts, the schema, or the model change.
The upsert derives the vacancy columns from the classification with these rules:
- One-to-one fields are columns:
title,job_function,role_category,seniority_level,references, andcompany_name_rawfrom the hiring organization.languageis the ISO 639-1 code of the vacancy text, or null for another value. - The model can miss a title that the page states. When the classification of
an open job page has no title,
titletakes the longest part of the page title, without the parts that name an organization of the classification or a label of the host name: "Bizzy - Internship" gives "Internship". A dash separates parts only between spaces. The stored classification keeps the model output, and the upsert logs each such title, so the rate is known. A generic page title, such as "Careers | Acme", gives a generic title. posted_dateandapplication_deadlineare the dates that the page states. A date without a year gets the year that puts it nearest toclassifiedAt, and a publication date is not later than it. A month without a day is the first day for a publication date and the last day for a deadline. A date without a month, or one that does not exist, is null. A rebuild from the stored classification gives the same dates.start_datefollows the same rule, with the first day for a month without a day.is_immediate_startis true for a start as soon as possible and null when no start is stated.residence_country_codesholds the countries that eligibility and remote-work conditions name. A region, such as the EU, has no code.- An array holds the distinct known values of all options:
engagement_arrangements,contract_terms, andwork_arrangement_modes. Null means that no value is known. No enum has anunknownvalue. - Workloads are alternatives.
minimum_hours_per_weekandmaximum_hours_per_weekcover all options, and so dominimum_days_per_weekandmaximum_days_per_week. A bound is null when an option does not state it in that unit. - Work arrangements are statements about one vacancy.
maximum_remote_days_per_weekis the highest stated value. - Pay comes from the first salary or contract rate:
minimum_pay,maximum_pay,pay_currency, andpay_period. - The hard filters of ADR-0014 come from requirements with level
requiredand no alternative.required_languagesholds the ISO 639-1 code of each such language requirement that names one language.required_education_levelandrequired_education_fieldshold the ISCED level and ISCED-F fields of the highest such degree, because a higher level meets a lower requirement. CEFR levels and the other requirements stay in the classification and the statements. - Each quote of the classification becomes a statement, in a fixed order: requirements, responsibilities, work conditions, work arrangements, engagements and their workloads, compensation, employment terms, benefits, and context. Contacts, organizations, and locations are not statements.
The database generates the search vectors. Each row indexes its own text:
vacancies.fts_vector the title (A) and the job function and company name (B),
and vacancy_statements.fts_vector the quote, with weight C for requirements
and responsibilities and D for the other types. A vacancy matches a search on
its own row or on any of its statements, and a statement match shows which
quote contains the term.
A vacancy with the same classification is unchanged, also when the page
markup changed. Then only last_seen_at changes, and the statement IDs stay the
same. A new classification replaces the statements.
Harvests and expiry
A harvest keeps one row over the retries of its job. The first attempt of a
configuration harvest job stores the harvest ID in the job data, and a retry
continues that row. One configuration runs at most one harvest: a harvest
start locks the configuration row, and a partial unique index allows one
running harvest per configuration. A running harvest that started more than
six hours ago stopped without completion, so the next start marks it failed.
Each source gets one ingestion per harvest, with the job ID
vacancy-<source ID>-<harvest ID>. When a retry finds other content for a
source that it already scheduled, the first ingestion stays.
A source expires when:
- Two consecutive complete harvests of its configuration do not list it.
missed_harvest_countcounts these harvests, and a listing resets it. An incomplete harvest expires nothing. A harvest is incomplete when it stops at a safety limit, or when it finds less than half of the previous item count of 20 or more. - Its page is a closed job, an error page, a job listing, or not a job page. A page that requires a login does not expire it.
A sitemap that cannot be read hides its sources, so it must not look like a
sitemap without jobs. When no sitemap gives URLs and one could not be read,
or the discovery timed out, the harvest fails and retries. When a sitemap that
a sitemap index or robots.txt lists cannot be read, the discovery times out, or
the limit of 25 sitemaps skips a listed one, the harvest is incomplete. A
failed guess, such as a missing /sitemap.xml, matters only when no sitemap
gives URLs.
A vacancy expires when all its sources have expired, and a closed job page closes it. When a harvest lists an expired source again, its ingestion makes the vacancy active.
Each upsert adds one to a count of its harvest: new_count, changed_count,
or unchanged_count for a job page, and rejected_count for another page.
The last failed attempt of a flow adds one to failed_count, and writes the
stage, the error message, and the time on the source (last_error_stage,
last_error_message, and last_error_at). The message of a pipeline error
starts with its code, such as [EXTRACTION_FAILED]. A later stored page
clears the error. The counts grow while the ingestions run, also after the
harvest row is complete. A retry of an upsert after its commit can count a
source twice. A job that fails because it stalled too often is not counted.
The classify, geocode, embed, and upsert jobs retry by error category: a
provider rate limit after 5 seconds, up to 5 times; a server error or timeout
after 2 seconds, up to 3 times; an unavailable geocoding service after 30
seconds, up to 3 times. A request that the provider rejects with another 4xx
status is not repeated. The harvest job keeps the fast default retries,
because it holds the admission of its host while it waits.
A site that answers 429 blocks its host for the wait of its Retry-After
header, at least 30 seconds and at most one day, and for 5 minutes without
the header.
Listing and sitemap URLs get one form: without fragment, credentials,
trailing slash, and utm_* parameters, and with sorted query parameters. A
page therefore has one source key in all strategies.
The pods of the vacancy workers give active jobs a termination grace period: 300 seconds for classify, 150 for embed, and 60 for the other stages. The shutdown timeout of each worker is 20 seconds shorter. A job that a deploy stops runs again, and the third stop fails it.
The database enums take their values from the tuples in
@gigflow/engine-jobs/vacancy/schema, and the Engine API builds its enums from
the database enums. The API list filters by seniority level, engagement
arrangement, contract term, work arrangement mode, the country of any location,
and a range of publication dates. The detail
returns the statements.
Alternatives
| Alternative | Reason for this decision |
|---|---|
| A table per classification array, such as engagements and compensations | Five more tables, a delete and insert per table on each change, and a join per filter. The filter columns on vacancies are faster for nearest-neighbor search with filters. |
| Only the JSON output, with JSONPath queries | Weak typing and queries that are hard to read. |
| Statements per revision (the earlier shape) | Every changed page hash added about 30 rows, also for the same content. |
| Statements on the classification | IDs that never change, but the vacancy is no longer the central record, and unused classifications keep their statements. An assessment of changed text is stale anyway. |
| A search vector that the upsert writes | The generated columns stay consistent without application code. |
| Store the Markdown | It is large, and the HTML in object storage gives it again. |
Exact checks of paired options, such as a freelance day rate for 5 days or an employment for 4 days, read the classification output in code. The SQL filters exclude a vacancy only on a conflict with the derived columns.
Consequences
Migration 0025 deletes all vacancies and source versions: they come from test
harvests and have no classification. Migration 0026 adds the seniority level,
the language, and the two dates, and replaces posted_at and closes_at,
which were always null. Migration 0027 adds the required languages and the
required degree, and migration 0028 the days per week, the start, and the
residence countries. Sources, configurations, and harvests stay, so the next
harvest classifies every source again. Pause vacancy harvests during the
deploy; flows of the old pipeline fail.
Migration 0029 adds the rejected count, the missed-harvest count, and the index of running harvests. Before it builds the index, it marks the older running harvests of a configuration failed.
Migration 0030 adds the error columns to vacancy_sources. Deploy the
vacancy workers with the flow producers: a flow job with the classified
backoff fails on a worker without that strategy. Sources whose URL form
changes get a new source once, and the old source expires after two complete
harvests.
An error page expires its source at once. A temporary error page therefore expires a vacancy until the next harvest lists the source and its ingestion accepts the page again.
A new derived column needs only code and a rebuild from
vacancy_classifications, without a model call.
Match assessments (ADR-0003) refer to statements with a cascading foreign key. When the classification of a vacancy changes, its statements and their assessments are replaced, and matching assesses the vacancy again.
Verification
Set TEST_DATABASE_URL to a migrated local database whose name contains
test. From the repository root, with Redis, PostgreSQL, and MinIO from
dev/cli/dev up engine, run each stage suite:
bun --env-file=apps/engine/apps/workers/vacancy/classify/.env \
test --timeout 30000 apps/engine/apps/workers/vacancy/classify/e2e
bun --env-file=apps/engine/apps/workers/vacancy/upsert/.env \
test --timeout 30000 apps/engine/apps/workers/vacancy/upsert/e2eThe classify test replaces only the model call. The upsert test runs the real
worker and database. Reports are written to
.context/verification/vacancy-classify/ and .context/verification/vacancy-upsert/.
The configuration harvest test runs the real worker, database, queues, and
storage, and replaces only the browser harvester. It checks the harvest retry,
the running-harvest lock, the expiry, and the item-drop guard. Its report is
written to .context/verification/configuration-harvest/:
bun --env-file=apps/engine/apps/workers/configuration/harvest/.env \
test --timeout 30000 apps/engine/apps/workers/configuration/harvest/e2eThe ingestion flow test runs the five vacancy workers for every HTML eval
fixture, from admission to upsert. It uses the local browser service, the real
model and embedding provider, and a Pelias fixture, because local development
has no geocoding service. It makes paid model calls, so it runs only when
VACANCY_FLOW_E2E is true:
VACANCY_FLOW_E2E=true bun --env-file=apps/engine/apps/workers/vacancy/upsert/.env \
test --timeout 3600000 \
apps/engine/apps/workers/vacancy/upsert/e2e/ingestion-flow.e2e.test.tsIts report, with the outcome, facts, and token usage of each page, is written
to .context/verification/vacancy-flow/.