Case 02
A job-hunting agent that cannot submit without my approval
A personal system that finds software and AI engineering jobs, scores them against my real background, drafts a tailored CV for each one I pick, and fills in the application, stopping whenever it needs my judgement. The foundations are built and deployed to a dev environment. The agent's behaviour is designed in full and is the next thing I build.
- Problem
- Job hunting is repetitive, but automating it risks sending a CV or an answer nobody approved.
- Built
- Foundations built and deployed to dev: database, sign-in, status rules. The agent that finds jobs and drafts CVs is designed, not built.
- Result
- The first roadmap phase is done and running in dev. No job has been found or applied for yet.
§1 Context & constraints
I wanted to automate the repetitive part of looking for work: finding roles, judging whether they fit, tailoring a CV and filling in forms. The part that is mine stays mine. The brief’s rule is to automate the work, not my judgement.
It is one workflow with three graphs. A scheduled discovery graph fetches jobs from several sources, normalises and deduplicates them, filters out clear mismatches with plain code, asks a model only for a triage call, scores each job with an explainable breakdown, and notifies me with numbered picks. When I say Apply #2, an application graph tailors a CV from my approved profile, checks it, waits for my approval, saves it to Drive, opens the form, fills what is already approved, waits for me on anything else, and submits only after a final approval. A third graph updates an application’s status from a plain sentence such as “I had the interview for #2”.
Around the graphs sit the parts that make it trustworthy. Every state change leaves an event in the same transaction, and that is built. The rest is designed: every model call will be recorded with its model, tokens and cost, with a budget checked before it is made, and an Agent Office, a small pixel-art office in the browser, will show what the backend is doing by reading that event stream and nothing else. The first version uses a language model for typed decisions. The second will swap a specialised decision model in, one decision at a time, and compare the two.
It is for one user first, me, with every table and query scoped to a user so it can serve more later. The documents are the source of truth: a project brief, an architecture with its data model and workflows, a tech stack that names the rejected alternative beside every choice, a roadmap, and a set of workflow diagrams.
Constraints
- Nothing is submitted without my explicit approval, and a tailored CV is never used until I approve it.
- The agent may not invent experience, skills, employers, dates or qualifications.
- CAPTCHAs and identity checks are never bypassed.
- Cost has to be visible and capped, with the budget checked before a model is called.
- Single-user first, but every table and query is scoped to a user so it can serve more people later.
- Simple first: one Postgres, with no separate queue broker or cache in the first version.
§2 Architecture
Quality over volume. Designed, not built. Jobs from several sources (Adzuna, Remotive, Arbeitnow and JSearch behind a flag) are normalised and deduplicated, and cheap deterministic filters reject clear mismatches before any model is called. A job that has not changed is never evaluated twice, and each score comes with a breakdown I can read.
A model only where judgement is needed. The worker skeleton runs in dev; the graphs are designed. Scoring, deduplication and budgets are plain code. Models handle triage, tailoring and classification at the cheapest tier that works, behind one decision interface, and a budget is checked before every call.
Real state, resumable. Built and deployed in dev. Workflow state, checkpoints and an append-only event log live in one Postgres, and the application's database role cannot update or delete the event tables. Every change writes its event in the same transaction, so the activity view and the notifications can only show work that happened.
Interrupts, not prompts. Designed, not built. Reviewing a tailored CV, answering a question the agent cannot answer truthfully, and the final review are all interrupts, answered from Telegram or the web. The run waits, survives a restart, and resumes when I answer.
One check guards submission. Designed, not built. The submit step raises unless a final approval and an approved CV version are on record. Before that, a validator blocks any CV claim with no matching fact in my profile. A CAPTCHA, an identity check, a video request or an assessment stops the run for me to handle, and none is bypassed.
The approved CV is the CV that is sent. Designed, not built. Drive holds the documents. The file id, revision id and checksum of the approved CV are pinned to the application. If Drive fails, the run stops instead of continuing on a cached copy.
§3 Decisions
Decision 1
Where approval lives
- Considered
- Asking the agent to confirm in its prompt, or trusting it to stop.
- Chose
- Every consequential step is a LangGraph interrupt, and the submit step asserts a recorded final approval and an approved CV version before it runs.
- Because
- A rule the agent is asked to follow can be missed. A check in code that raises cannot. It also makes the workflow resumable, so I can answer hours later.
- Trade-off
- A run can wait for me indefinitely, so waiting needs its own state and its own notifications.
Decision 2
How a CV is tailored without being falsified
- Considered
- Letting a model rewrite the CV freely.
- Chose
- A tailored CV is structured data in the shape of my base profile. It can reorder and emphasise, but it cannot add an employer, a date, a degree or a skill, and every claim is checked against a fact in the base profile.
- Because
- Truth is the one thing a job application cannot recover from. A structure a model has to fill is easier to check than prose.
- Trade-off
- The CV can only be as good as the base profile, and the check may reject honest rewording. How strict the matching can be is still to be tested on real CVs.
Decision 3
What counts as the record
- Considered
- Keeping the CVs in the application's own database, with a history that the application can rewrite.
- Chose
- Google Drive holds the documents, and Postgres pins the file id, revision id and checksum of the exact CV used for each application. The history is append-only: the application's database role has no UPDATE or DELETE on the event, decision and call tables.
- Because
- The CV stays mine, in a place I already use, and I can prove which version was sent. A history the application cannot edit is one I can trust. The grants are in the baseline migration, applied in dev.
- Trade-off
- A Drive outage stops an application instead of letting it continue on a cached copy, and fixing a wrong row means the owner role, not the application.
Decision 4
How to learn what a specialised model could do
- Considered
- Building the whole system around a specialised decision model from the start.
- Chose
- Every typed decision goes through one interface, so a decision point can be switched from a language model to a specialised one, run in shadow mode, and compared on accuracy, correction rate, latency and cost.
- Because
- I want to replace one decision at a time and measure it, not adopt something because it is new.
- Trade-off
- The interface, and the logging of every decision and every correction I make, have to be built before there is a second engine to use them.
§4 The hard part
The hard part, so far, is making the safe path the only path. An agent that can fill in and submit a form for me can also send something false, or something I never saw, and a note in a prompt asking it not to is not a control.
So the design starts from what must be impossible, and the foundations already enforce the part that does not need the agent. The application’s database role cannot update or delete the event tables. Every state change writes an outbox event in the same transaction, so the activity feed I will watch can only show work that really happened, and a change that is rolled back announces nothing. The application’s fifteen statuses and their allowed moves are a whitelist in the domain package, and anything else is rejected.
The rest is planned, and I have not proved it yet. The submit step must raise unless a final approval and an approved CV are on record, with a dedicated test for exactly that check. And a tailored CV must be blocked when any claim in it has no matching fact in my base profile. What I do not know yet is how strict that check can be before it rejects honest rewording. That needs real CVs and real job posts, and it is the first thing the agent phases have to answer.
§5 Outcome
What is built, and running in a dev environment, is the foundation: the first phase of the roadmap. A monorepo that runs the whole stack locally with one command. The domain models with the status whitelist. The database, with every table, the append-only grants and the outbox trigger. Google sign-in with signed tokens, and a test that one user cannot read another’s rows. A Terraform baseline for the cloud. And CI/CD that tests every change and deploys the api to a dev environment, with the web app on Vercel.
Nothing about jobs exists yet. Onboarding, discovery, CV tailoring and checking, the application flow, Telegram, the Agent Office and the comparison with a specialised decision model are designed and not built. No job has been found or applied for.
I will update this page as each phase lands.
What I'd do differently
- I wrote the whole design, then built the foundations, before the agent touched a single job. Next time I would build the thinnest end-to-end run first, one job from discovery to a reviewed application, so the design meets real data early.