Gigflow docs
Engineering

ADR-0004: Source-validated vacancy ingestion

Shared source preparation, classification, and passage checks.

Status: Proposed. Implementation available for review.

Date: 2026-09-21

Decision

Use one source preparation path for HTML evals and stored documents. The document parser and parseHtmlMarkdown share HTML preparation and media exclusions. Retain semantic regions unless a saved noise filter approves removal. Decode explicitly marked rich-text containers before cleaning. Do not interpret arbitrary escaped code examples as HTML.

The pipeline is harvest, preprocess, classify, enrich, embed, and finalize. Preprocess prepares text without an LLM call. Classify uses the three-call classifier, with fresh messages and the same Markdown for each call. Code checks source passages before returning the classification. Missing passages stop the flow with a non-retryable extraction failure. An unknown page also stops the flow. Neither failure means that a candidate does not match.

The classification schema lives in @gigflow/engine-jobs/vacancy/schema. vacancy-content imports it from there. This avoids a dependency cycle between the shared worker contracts and the classification capability.

Save the result

Save the complete current classification in vacancies.content. Increment its revision when accepted content changes. Save the exact source Markdown with the source version and its raw artifact reference. This is not a separate table of classification audit records.

Create statement rows from source quotes in code. Preserve the type, requirement level, source version, revision, and path into the classification. Keep previous statement revisions so their IDs remain valid. Readers must select statements for the vacancy's current revision unless they request history.

Save content, statements, locations, and embeddings in one transaction. Check quotes again at the internal finalize API before the transaction. Clear old claims on an accepted update. The claims queue is absent from new flows; its worker rejects jobs that still use the retired stage.

Flat legacy columns are display summaries, not the complete matching model. Do not collapse alternative contracts or scoped pay into a single hard filter. Those flat fields remain null when the migration cannot represent their scope. Contract options, hours, pay, and remote limits remain in the structured content. This change does not populate vacancy_terms or enable automatic hard filters.

Use the shared search-document builder for two broad embedding inputs. Preserve the configured embedding provider and dimensions in this migration. Store the work text vector in the existing context vector column and clear the unused responsibility and culture vectors. A provider or dimension change requires coordinated candidate and vacancy re-embedding. The old claims-based matching scorer is not migrated by this ingestion change.

Deploy and verify

Apply migration 0017_vacancy_source_validation before starting the changed workers and Engine API. Stop or drain old flows first, then create new runs. Old preprocess and classify payloads are not compatible with these contracts. No queued jobs or deployed services are changed by editing this code.

Run passage tests, worker contract tests, database integration tests, and the split model evals. A quote match proves text presence only. It does not prove correct interpretation, correct scope, or complete extraction. Critical eval failures must still prevent release of the affected automatic filters.

On this page