ADR-0005: Vacancy HTML noise filter
Review the regions that HTML cleanup would remove and retain uncertain content.
Status: Proposed
Date: 2026-09-21
Context
A vacancy page can have a site footer and an application footer. Both use
the same HTML tag. Navigation and other regions removed by onlyMainContent
can also contain job facts. Removing all such regions can remove useful data.
Decision
DocumentClient selects the format parser in @gigflow/harvest/documents.
Each parser returns Markdown through the same result contract. HtmlFormat.parse
owns HTML preparation and includes its preparation report in that result.
NoiseFilter.create learns rules in a separate configuration worker.
applyNoiseFilter checks those rules during preparation. RegionClassifier
owns the model call, with a separate prompt and Zod schema.
getCurrentNoiseFilter checks the stored filter size, schema, and configuration
hash. Preparation and review each use this function once. The learner receives
the validated filter or null.
Only regions that the existing footer, nav, and onlyMainContent cleanup
would remove are review candidates. Ordinary page content is not classified.
Review a complete outer region once. Keep mixed content and uncertain regions.
The model returns supplied region IDs, decisions, and reasons. It cannot write
selectors. There is no model call in vacancy preprocessing.
Preparation removes assets first. It resolves links against the source URL and
any usable base URL before removing the document head. For an href listing,
a selector alone does not authorize removal. Each matched card must contain links
to other jobs that pass the strategy's job URL and exclusion patterns. Keep links
to the current job, including fragment links. Keep containers with text outside
those links, an h1, or form and table content. Missing or invalid URL patterns,
invalid selectors, and selectors that match document roots retain content.
Sitemap strategies have no card selector.
For a click listing, the configured item selector authorizes removal without
links or a URL pattern. Keep a match that contains an h1, is the configured
detail region, or contains that region. If the clicked item itself supplies the
detail content, keep its content. Click harvesting applies the same removal
before it converts each detail to Markdown, including pages that keep their URL.
Invalid selectors and document roots retain content.
Stored text and other document formats keep their existing parsers. Existing raw artifacts and their hashes do not change. Future click captures contain the prepared detail text.
The learner reads 3 to 6 distinct, usable HTML source versions. The API selects the latest version of each active source before it checks the page type. Each sample must belong to the active configuration and company. The worker checks the raw bytes against the stored hash. Failed reads do not count as samples. The worker reads only current samples. The learner reuses approved removal rules when their selectors and signatures still match. It reviews other eligible regions again, including regions retained by a previous review. This can increase model calls when the sample set changes. Manual review uses no previous decisions.
Source and scheduled requests compare the selected usable samples with the saved samples after file validation. If the sets match, they skip learning and do not consume a review slot. Skipped references alone do not start another review. Manual and changed-content requests still review unchanged samples. The worker compares references before and after learning to report changes during the review.
Every region sent to the model must receive exactly one decision. Missing, duplicate, or unknown region IDs fail the batch. If any model batch fails, the worker retries the review without saving partial results. A complete review with no removal rules remains valid.
Learned selectors use tags, roles, labels, and direct ancestry. Each selector must match one complete region in every training page. The complete region must have the same signature across those pages. The signature includes text, element structure, and attributes other than IDs, classes, and inline styles. It includes absolute links. Regions with controls that Markdown cannot fully represent, ambiguous selectors, different signatures, or oversized input remain in the page. At application time, changed or ambiguous regions remain and request review.
Storage and worker lifecycle
Store the optional noiseFilter object in the existing configuration JSONB.
It contains a version, configuration hash, removal rules, sample references,
and completion time. No database migration or extra state table is required.
The configuration hash includes the canonical URL, schema version, and strategies.
| Stored state | HTML preparation |
|---|---|
| Filter absent | Retain semantic regions and request background review |
| Valid filter with no rules | Retain all semantic regions |
| Valid filter with rules | Remove only matching approved regions |
| Invalid, stale, or unreadable filter | Retain semantic regions and request review |
Internal admin tRPC procedures own sample selection, paged reconciliation, save, and reset. Each endpoint keeps its schema and SQL in its own file. Save locks the configuration and sample sources. It checks ownership and latest source versions again. A configuration row timestamp prevents an old completion from overwriting a reset or another completion. Reset advances this token even when no filter exists. Identical rediscovery preserves the filter. Changed discovery configuration removes it.
Vacancy finalization requests review after commit, including unchanged valid sources. An enqueue failure does not fail finalization. A stable cursor scan of active configurations runs every 15 minutes to recover missing requests. Automatic requests share one job ID per configuration. Manual requests use a separate ID so an automatic delay cannot suppress them. Terminal jobs are removed so later reviews can use those IDs again. Initial global queue concurrency is one.
Automatic review cycles are limited to one per hour per configuration. A cycle permits three attempts for transient failure. Manual review bypasses the hourly limit. Shared queue deferrals use steps of at most 10 minutes. A source committed during an active review is eligible in the next reconciliation scan.
Limits
| Resource | Limit |
|---|---|
| Candidate source references | 20 per scan |
| Usable current HTML samples | 6 |
| File bytes / total bytes read | 8 MiB / 24 MiB |
| Elements per document | 100,000 |
| Candidate regions | 32 |
| Regions / prompt bytes per request | 4 / 24 KiB |
| Model calls per attempt | 8 |
| File read / model call / job deadline | 20 seconds / 30 seconds / 5 minutes |
| Saved rules / selector length / filter size | 16 / 512 characters / 32 KiB |
Functional validation
HTML preparation and background learning are always active. There are no rollout modes or environment switches. Until a valid filter exists, preparation retains semantic regions and removes only assets and configured job cards. The local Engine profile and image build discover the worker package automatically.
Use these structured log events to check operation:
| Event | Check |
|---|---|
Noise filter review skipped | insufficient_samples means fewer than three usable pages; unchanged means the saved review is current |
Noise filter regions reviewed | Check status, modelRequests, failedRequests, and skipped counts |
Noise filter review finished | saved confirms persistence; superseded means the configuration or evidence changed |
Vacancy HTML prepared | Check the configuration ID, path, appliedRules, skippedRules, and needsReview |
A zero-rule filter is valid when all regions must remain. Validate real model decisions through the complete learning and preparation path. Compare salary, benefits, location, requirements, application links, deadlines, and employer text. The prepared Markdown and extracted fields must retain those facts.
Tests cover separate footers, navigation, mixed parents, changed links and
benefits, invalid selectors, partial model failures, invalid model decisions,
repeated reviews, missing files, queue delays, and nonfatal enqueue failures.
PostgreSQL integration tests cover conflicting
saves, reset races, ownership, latest-version selection, and rediscovery.
Run the database tests against a disposable database whose name includes _test_.
bun run --filter=@gigflow/harvest test
bun run --filter=@gigflow/engine-workers-configuration-noise-filter test
ENGINE_DATABASE_TEST_URL='<disposable database URL>' bun test --isolate apps/engine/apps/api/src/__tests__/trpc/noise-filter.e2e.test.ts