Architecture

One process, four agents, one datastore.

There are no microservices, no queue and no cache tier. The workload is one CPU-bound request at a time against an 800-role corpus held in memory, and every boundary below exists because something needed it — not because a diagram looked better with it.

System diagram

Everything inside the dashed boundary is a single Python process with one worker. Two workers would double the 215 MB resident set to serve a queue of one.

Browser · Next.js 16 on Vercel 8 routes · NEXT_PUBLIC_API_URL inlined at build HTTPS FASTAPI EDGE CORS allowlist Rate limit · per client IP Upload guard · magic bytes one process · one uvicorn worker · 215 MB resident AGENT 1 Parser pdf · docx · txt → text AGENT 2 Extractor skills · years · degree AGENT 3 Scorer 5 components + ML AGENT 4 Explainer top-K only, never 800 Vocabulary 679 canonical skills 1,554 aliases · O(1) also read by Agent 3 Model artifacts joblib · 1 predict_proba for the whole corpus not one call per role LLMProvider + budget ollama · openrouter rule_based (fallback) quota · semaphore · K ≤ 3 Document readers PyMuPDF · pdfminer python-docx magic bytes, then parse SQLITE · WAL · ONE TRANSACTION PER UPLOAD jobs match_history llm_quota seeded once from data/json/jobs.json · reseeding an already-populated table is a no-op the JSON file is a seed, not a source of truth — POST /jobs writes here

Request lifecycle

What one upload actually does

StepWhereDetail
1 Guardsrc/api.pyOrigin checked against an allowlist, per-IP limit charged, real file bytes checked against the extension, 10 MB cap
2 Parseagent1_parser.pyPDF, DOCX or TXT to a clean text layer. Under 50 extracted characters is a parse failure, not a weak candidate
3 Extractagent2_extractor.pySkills, years, degree and contact details; every skill resolved through the alias index
4 Scoreagent3_scorer.pyAll 800 roles in one pass — five rule components per role, then one vectorised model call for the corpus
5 Rankpipeline.pySort, slice to top-K. Explanations are generated after ranking, so the LLM never sees 800 roles
6 Explainexplaining/At most 3 calls, budget checked first, fallback per item rather than per batch
7 Persiststorage/database.pyTop-K written in one transaction, not one connection per row

Design decisions

Six that needed an argument

Each replaced something that worked. The reasoning is the interesting part, and each one is checkable against a named file.

The alias index is built once, at load

Every language normalised to its category name. _get_canonical_skill read a category-nested dictionary as if it were flat, so unrelated skills matched at 100%. Skills are 50% of the rule score, which is 60% of the hybrid score. One alias→canonical map at load makes lookup O(1).

['programming_languages', 'devops'] → ['FastAPI'], and a backend role replaced a frontend one at the top.
src/core/vocabulary.py · load_alias_index()

The provider is fixed at construction

A singleton reassigned per request, restored outside a finally. A lock would have made the symptom rarer; injecting once removes the mutable state. A typing.Protocol with two methods, so implementations relate by shape and a fake is about five lines.

Three real implementations. No abstraction arrives here without two.
src/agents/explaining/__init__.py · build_provider()

One model call for the corpus, not one per job

The classifier ran 800 times per upload. Each call rebuilt a one-row DataFrame and re-ran the fitted transform, so most of the request was framework overhead. One frame, one transform, one predict_proba.

8.13 s → 0.033 s, with a guard test asserting identical scores.
src/agents/scoring/ml_scorer.py

Persist the answer, not the working

Every upload wrote ~4,000 rows and returned ten. The database was recording the search space rather than the result. Now: sort, slice to top-K, one executemany transaction with WAL.

114× cheaper per row. The test counts connections rather than milliseconds — a fixed budget measures the machine.
src/storage/database.py · save_matches_batch()

Explanations are bounded three ways

An unmetered outbound bill, triggered by one upload. Explanations ran for every role scoring ≥ 0.6, bounded only by corpus size — mostly for roles that never made the top ten. Now: rank first, cap at three, a daily quota in SQLite, and a concurrency semaphore.

Degrades rather than breaks. At 90% of quota it switches to rule-based, before the provider starts refusing.
src/agents/explaining/budget.py · pipeline.py

A file's bytes, not its extension

A .exe renamed .pdf reached the parser. Validation checked the filename, drag-and-drop silently discarded DOCX and TXT, and the interface advertised a limit the API never enforced. Leading bytes are now checked before any parser runs.

10 MB, checked before read. Reading first and rejecting after is a limit you do not have.
src/api.py · read_upload()

Deliberately absent. No microservices, no Celery/Redis queue, no vector database, no Kubernetes, no GraphQL, no repository pattern over SQLite. Semantic skill matching via embeddings is the one genuinely tempting exclusion — it would improve scoring without putting a model in the hot path — but it needs a vector store, an embedding model and an index build, and it was deferred rather than quietly skipped.