ADR-0003: Evidence-based vacancy matching
A proposed plan for retrieval, hard filters, and match explanations.
Status: Proposed. No implementation included.
Date: 2026-09-17
Decision
Use Qwen3-Embedding-8B through Upgreat for retrieval. Use source evidence, structured checks, and an LLM for explanations. Similarity is not proof of ability or a match percentage.
Keep skills and domains as open text, without a fixed taxonomy. Replace flat vacancy claims with evidence records.
Build the data foundation
Use these records. Names are proposals, except for vacancy_embeddings.
| Record | What to store | Why |
|---|---|---|
| Classification version | Full output, source version, schema version, and extraction model | Preserve the input to every derived record. |
vacancy_evidence | ID, vacancy ID, classification version, kind, type, requirement level, source quote, and source path | Explain each requirement without losing its meaning. |
| Structured filter facts | Typed attendance, location, contract, rate, and availability facts, with evidence IDs and applicable contract option | Check explicit conflicts with SQL or code. |
vacancy_embeddings | Separate broad vectors for requirements and for work/context | Find relevant vacancies across millions of records. |
| Candidate evidence | Each experience or project, its activities, dates, employer, end client, and source | Support detailed comparisons, including domain experience. |
| Match assessments | Criterion ID, candidate evidence IDs, result, explanation, and input/model versions | Give both sides consistent, traceable explanations. |
Map classification arrays to evidence rows in code. Keep “React or Angular” together. Verify quotes and preserve evidence IDs within each version.
Use an extra LLM call to normalize filter facts where needed, once per extraction version. Validate its output; never generate SQL. Exclude only established conflicts. Keep unknown values and contract alternatives. “Required” does not automatically mean “safe SQL filter.”
Keep capability separate from preference
Separate demonstrated capability from desired work. Match capability against requirements and preferences against work/conditions. Interest is not experience.
Store complete experiences. Add experience embeddings for better discovery or
evidence selection if tests support this. Keep vacancy_evidence_embeddings
optional and separate. Embed meaningful statements, not every sentence. Deep
assessment reads the original evidence.
Configure Qwen and search
Use qwen3-embedding-8b. Confirm access, dimension controls, batching limits, and
data handling with Upgreat.
Qwen3-Embedding-8B has 4,096 full dimensions, not 2,560. Put English task instructions on queries, normally not documents. Index vacancies for candidate searches and candidates for vacancy searches. Change query instructions with direction. See the Qwen model card.
Test 1,024 dimensions against a full-dimensional reference. Standard pgvector
HNSW indexes support 2,000 dimensions for vector and 4,000 for halfvec.
Full 4,096-dimensional vectors need another indexing strategy. See
pgvector indexing limits.
Version vectors by model, dimensions, input hash, purpose, and instruction template. Re-embed both sides when switching from OpenAI; do not mix models.
Retrieve, assess, and explain
Use these counts as initial test settings, not fixed quality guarantees.
- Combine reliable filters, approximate nearest-neighbor search, and PostgreSQL full-text search. Preserve exact technical terms. Why: meaning alone can miss specific products.
- Merge results to about 300 unique pairs using reciprocal-rank fusion. Cap experience-query contributions. Why: longer careers must not automatically receive higher scores.
- Rerank the actual pair texts to about 50. Evaluate Qwen3-Reranker-8B as a separate service choice. Upgreat's public API does not offer reranking. See Upgreat API differences.
- Assess the first 20–50 pairs with an LLM, criterion by criterion. Return
met,not_met,unknown, ornot_applicable, with evidence references. Why: ranking scores do not explain requirement coverage. - Check dates, non-overlapping durations, rates, and attendance with code. Explain both sides from the saved assessment. Why: explanations must agree. Missing evidence is not proof of missing ability.
Payments projects can support domain experience; a bank employer alone cannot. React and AngularJS indicate related work, not equivalent experience.
Validate before scaling
Measure shortlist recall, relevance, false exclusions, citation accuracy, cost, and latency. Use reviewed Dutch, French, and English cases, including domain experience, framework differences, unknowns, and contract alternatives. Test at the expected database size.
Build evidence, filter facts, and broad vectors first. Add granular vectors when tests show a gain. Use Engine API/SDK contracts and authorized candidate access.
okay. I agree. if we look at the current classify vacancy extraction. I need to have a new table vacancy-evidences where i store
- still remain vacancy_embeddings and make those pretty broad.
- potentially have embeddings on vacancy-evidences?
Do i got that right? or what else should happen? Yes. Keep broad vacancy embeddings, add evidence records, and make evidence embeddings optional. But you also need structured facts for hard filters.
These have different purposes:
Data Purpose Classification output Preserve the complete extracted vacancy vacancy_evidence Preserve individual requirements, conditions, and context, with their source text Structured filter facts Check clear constraints with SQL vacancy_embeddings Find potentially relevant vacancies quickly vacancy_evidence_embeddings — optional Find specific requirements or context that broad search can miss This is the design I suggest.
- Store evidence from the current classification
Your current requirements schema:291 already returns type, level, and quote. That is a useful starting point.
For example, the classifier could produce these evidence records:
ID Kind Type Level Source quote E1 Requirement Technology Required “Experience with React or Angular is required.” E2 Requirement Domain Preferred “Experience with payment processing is an advantage.” E3 Work condition Office attendance — “You must attend our Brussels office at least two days per week.” E4 Context — — “The team builds payment reconciliation software for merchants.” Each record should contain at least:
id and vacancyId extractionId, linked to the saved classification and source version kind: requirement, responsibility, work condition, context, etc. type and level, where applicable quote: the original source text sourcePath: for example, requirements[0] One row should contain one complete statement, not necessarily one sentence. Keep “React or Angular” together. Otherwise, it can become two mandatory requirements by mistake.
Creating these rows from the classification does not need another LLM call. Code can do it. Also check that the quotes occur in the source.
These categories describe the structure of the data. They do not require a fixed skills or industry taxonomy.
- Create structured filter facts separately
An evidence record preserves what the vacancy says. A filter fact gives code a value it can compare.
For E3, the normalized facts could be:
{ "evidenceIds": ["E3"], "officeLocation": "Brussels", "minimumOfficeDaysPerWeek": 2 } If the candidate explicitly accepts zero office days, code can identify a conflict.
If the candidate has not specified this, the result is unknown, not “does not match.”
With the current quote-based schema, some facts need an interpretation step. For example, attendance rules and contract-specific conditions. A focused LLM normalization call can produce these facts. Later, you could include them in classification if testing shows that this is reliable.
The model produces validated data. It does not generate SQL.
Also, “required” does not automatically mean “SQL hard filter.” A required payments background can remain a criterion for deep assessment when you cannot check it reliably through structured data.
- Keep broad embeddings, but give each a clear purpose
“Broad” should mean a useful summary, not all vacancy text mixed together.
For an initial design, I would test:
Work and context: responsibilities, product, project, and domain. Candidate requirements: skills, experience, and qualifications. These embeddings help find vacancies. They do not prove a match.
Save the embedding model, input hash, and extraction version. This prevents old vectors from being used as if they describe a changed vacancy.
- Make evidence embeddings an optional search index
I would use a separate vacancy_evidence_embeddings table if you add them.
Each vector points to an evidence record. You can then change the embedding model or rebuild vectors without changing the source evidence.
For example, an embedding for E4 could help retrieve this vacancy for someone with payment reconciliation experience, even if the broad vector mainly reflects frontend development.
I would not embed every evidence row at the start. Start with broad retrieval plus keyword search. Test whether it misses relevant vacancies. Add requirement or context vectors where they improve retrieval.
- Deep matching still uses the evidence text
After retrieval, load the vacancy evidence and the candidate’s relevant experience records. The detailed assessment connects them:
E2 is supported by candidate experience C7: the candidate implemented payment settlement and reconciliation in a merchant platform.
Save that assessment with both evidence IDs, its result, and its explanation.
It can also say:
E1 has related evidence, but direct React or Angular experience is not established.
The evidence vectors are not needed to make either assessment. The model reads the actual text.
For domain matching, preserve which employer, end client, project, or product each experience describes. Employment at a payments company alone does not prove payments work.
My recommendation: build versioned classification storage, evidence rows, structured filter facts, and broad embeddings first. Keep evidence embeddings optional. The foundation for deep matching is traceable evidence—not a vector for every sentence.
No code changed.