ADR-0010: Vacancy classify stage
Prepare vacancy Markdown and classify it in one job.
Status: Proposed. Implementation available for review.
Date: 2026-09-23
Context
ADR-0007 keeps the vacancy ingestion sequence
harvest -> preprocess -> classify -> enrich -> embed -> finalize. Preprocess
reads the harvested file, prepares Markdown with the configuration's noise
filter (ADR-0005), and requests a noise-filter review when the filter needs one.
It makes no model call. Classify receives only that Markdown. The separate
stage adds a queue, a worker deployment, and a stored intermediate result, but
no independent retry or capacity boundary: preprocessing is cheap and
deterministic.
Decision
This ADR replaces the preprocess stage in the ADR-0007 sequence. The vacancy ingestion sequence is:
harvest -> classify -> enrich -> embed -> finalizeThe classify job reads the harvested file from object storage. It loads the
source's configuration from the Engine database and uses it only when it is
active. DocumentClient.parseDocument prepares the Markdown. When the
preparation report needs a review, the job enqueues a changed noise-filter
review for that configuration. Then the three-call classifier runs on the
prepared Markdown.
The steps before the model calls do not ignore failures. A database read or a
review submission that fails fails the job, and the job retries. A retry repeats
only the document preparation before it calls the model again. A non-retryable
HarvestError, such as an empty or unsupported document, fails without a
retry.
The classify result keeps the harvest, identity, preprocess.markdown, and
classification fields, so enrich and embed accept it unchanged.
Consequences
The vacancy-preprocess queue, job contract, and worker are removed. Flows
published with the old graph wait for a preprocess job that no worker runs.
Each keeps an admission slot, because the scheduler frees a slot only when the
flow root completes or fails. Drain the queue before deployment, or cancel those
ingestions afterwards with VacancyClient.cancelIngestion, which removes the
flow and frees the slot. Do not delete their jobs directly: a request whose flow
root was deleted keeps its slot.
Producers that still build the six-stage graph through
@gigflow/engine-job-client fail in the vacancy-harvest worker before this
change. They must move to VacancyClient.enqueueIngestion.
Classify now needs database, object storage, and noise-filter queue access.
Verification
Set TEST_DATABASE_URL to a migrated local database whose name contains
test. From the repository root, run:
bun --env-file=apps/engine/apps/workers/vacancy/classify/.env \
test --timeout 30000 apps/engine/apps/workers/vacancy/classify/e2eThe test runs the real worker with Redis, PostgreSQL, and MinIO from
dev/cli/dev up engine. A preload replaces only the model call. The report is
written to .context/verification/vacancy-classify/worker-new/report.json.