← Personal projects Applied AI Personal project · 2026

An LLM pipeline where the model goes last

A local-first job-search tracker: it captures a posting, extracts the facts, scores a résumé against it and logs every status change. The design rule is deterministic first, model last — a local LLM only fills the blanks that APIs and structured data left empty.

  • Python
  • FastAPI
  • SQLite
  • Ollama
  • React
  • Chrome MV3
3
Parsers tried before the LLM
6 GB VRAM
Runs in
65% keywords + 35% embeddings
Fit score
46
Backend tests

Architecture

  1. 01CaptureURL or browser extension
  2. 02ATS APIexact JSON
  3. 03JSON-LDstructured data
  4. 04Readabilitymain-content extract
  5. 05Local LLMschema-constrained, fills blanks
  6. 06Event logSQLite, append-only

The problem

A job search generates a surprising amount of data and almost none of it is structured. A posting is a web page. The facts you need from it — required skills, seniority, salary, whether they sponsor — are buried in prose. And the question you actually care about three weeks later, which of the things I’m doing is working, can’t be answered from a spreadsheet of company names.

The obvious 2026 answer is to hand the page to a language model and ask for JSON. That works until it doesn’t: the model invents a salary, misreads a location, and you can’t tell which fields to trust. So this project is built around the opposite rule.

Deterministic first, model last

Ingestion is a chain of parsers, ordered from most exact to least:

  1. ATS APIs. Greenhouse, Lever, Ashby and Workday expose public JSON for their postings. If the URL matches one, the API is called directly. Clean title, company, location, description and salary, with no scraping and no guessing.
  2. JSON-LD. Search engines require JobPosting structured data, so most career sites embed it. Exact fields again.
  3. Readability extraction. For arbitrary HTML, the main content is pulled out and the title and company come from the page’s metadata.
  4. The LLM, as a background task. A small local model through Ollama, given a JSON Schema as its output format so it can only emit the typed record: skills, tools, seniority, years of experience, salary, a two-sentence summary.

The rule that makes this trustworthy is in step four: the model only fills blanks. It never overwrites a value that steps one to three produced. If the ATS API said the job is in Tempe, no model gets a vote.

That ordering also makes it cheap. Most postings resolve at step one or two and the model is only asked for the fields nothing else could provide. The whole thing runs on a laptop GPU with 6 GB of VRAM.

The browser extension solves a data-access problem

Some postings sit behind a login, and fetching the URL from a server gets you a sign-in page. The Chrome extension sidesteps that by sending the DOM the browser is already rendering, so the backend never makes the request at all.

The event log is the point

Each application has a status column, and that column is a convenience for filtering. The truth is a separate append-only table: every status change, with a timestamp, never updated or deleted.

That one decision is what makes the analytics possible. A funnel — how many reached a phone screen, how many reached an interview — is a query over events. So is median time to first response, and so is the flag for applications that have gone quiet. None of those can be computed from a table that only remembers the current state.

Résumés are stored by the SHA-256 of the file, so the same PDF can’t be saved twice under two names and every application records exactly which bytes were sent.

Scoring a résumé against a posting

Two signals, combined:

  • Keyword coverage, 65%. Which of the posting’s required skills appear in the résumé. It’s blunt, and it’s the part you can act on: the missing skills are listed.
  • Embedding similarity, 35%. Cosine similarity between embeddings of the résumé and the posting, to catch overlap that exact keywords miss.

The weights lean toward the transparent signal on purpose. A score nobody can explain is not much use for deciding what to change. If the embedding model isn’t running, the score falls back to keywords alone rather than failing.

What broke

A stale server served HTML to the API client. The single-page app’s fallback route answered any unknown path with index.html, including API routes that only existed in newer code. An out-of-date server therefore returned a web page where the UI expected JSON, and the page crashed on undefined.length. The fix was three small things: unknown API paths now return a JSON 404, the client rejects non-JSON responses with a clear error, and the UI shows that message instead of a blank screen.

React forms ignore el.value = x. React keeps its own copy of an input’s value, so setting it directly changes what you see and not what gets submitted. The extension has to call the prototype’s value setter and then dispatch the events React listens for.

What I’d do differently

Evaluate the extraction. The schema guarantees the output is well-formed, not that it’s right. A labelled set of postings with known skills and seniority would give a real accuracy number per field, and tell me whether a larger model is worth the VRAM.

Calibrate the fit score against outcomes. The event log already records which applications got a response. The 65/35 split is a judgement call that the data could now test.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.