Four-agent CV screening pipeline

Every score, broken into the parts that made it.

Upload a résumé. It is parsed, its skills resolved against a controlled vocabulary of 679 canonical terms, and scored against all 800 roles in the corpus in a single pass — with five weighted components returned separately, so any total can be reconstructed from its parts rather than taken on trust.

The live API runs on a development machine behind a tunnel, so the demo is only up while that machine is. If it does not respond, the walkthrough is a full capture of the same app running.

800roles in the corpus
679canonical skills
544tests passing
84.07%branch coverage
15API operations

What this proves

Six problems that only show up when you build the whole thing

Each was found here, reproduced, and fixed with a test that failed first. Every card names the file, so the claim can be checked.

Concurrency

A race that only exists above one worker. Choosing a provider per request meant writing to a module-level singleton and restoring it outside a finally. The fix was not a lock — the provider is injected once at construction, so there is no mutable state to contend for.

Structurally impossible, rather than defended against.
src/agents/explaining/__init__.py

Determinism

The same input scored two different numbers. Keyword extraction sliced [:20] from a set, and Python randomises string hashing per process. Terms now rank by frequency, tied alphabetically.

0.40 or 0.45 for one CV and job, across five processes.
src/agents/agent3_scorer.py

Idempotency

Seeding that cannot overwrite your data. The corpus became writable through POST /jobs, and on an ephemeral filesystem the obvious implementation reseeds every deploy. seed_jobs writes nothing into a populated table.

A seed, not a source of truth. Redeploy is a no-op.
src/storage/database.py

Trust boundary

Rate limits charged to the right client. Per-IP limits read the socket address, which behind a proxy is the proxy's — so every visitor shares one bucket. X-Forwarded-For is client-supplied, so it is believed only when TRUST_PROXY_HEADERS says something is in front.

Off by default. Trusting it with nothing in front is the worse mistake.
src/api.py · client_address()

Honest metrics

Target leakage found in my own model. A perfect ROC-AUC. Removing the feature I suspected changed nothing. The label turned out to be a threshold on a column already excluded from training, and two ordinary columns rebuild it anyway.

No accuracy figure is published, and a contract test asserts none appears.
models/production/model_metadata.json

Silent failure

A degraded run must not look like a healthy one. A rule-based explanation and a model-written one are both fluent paragraphs, so a dead key produces a broken demo that reads as a working one.

Three signals: explanation_source, scoring_mode, and the provider named at /health.
src/agents/explaining/protocol.py · src/api.py

Built with

Stack

Versions as pinned in requirements.txt and frontend/package.json.

Python 3.10+ FastAPI 0.104.1 Pydantic 2.12.5 scikit-learn 1.3+ SQLite WAL slowapi rate limits Next.js 16.3.0 React 18.3.1 TypeScript 5.4.2 Tailwind CSS 3.4.1 pytest 8.4.1 mypy 1.19.0 ruff 0.16.2 black 25.11.0

The lint toolchain is pinned exactly, not by lower bound. A range makes "correctly formatted" a function of the install date rather than of the code: an open black>=24.0.0 had CI resolve a major version ahead of a developer's machine, and the two disagreed about seven files nobody had touched.

How a score is built

Five weighted components, each returned separately

ComponentWeightWhat it measures
Skill match50%CV skills against required and preferred, both resolved to the vocabulary
Experience20%Years against the stated range, penalising both directions
Title similarity17%The candidate's role against the job title
Education8%Highest degree against the stated requirement
Keyword overlap5%Terms from the description present in the CV

Enforced, not documented. The weights live in config/agents.yaml and nowhere else. A set that does not sum to 1.0 fails at import with the offending values named, and a contract test asserts the five components add up to the reported rule-based total — because a stacked bar built on drifting weights does not look wrong, it looks right.