Why I built it
A job search is three jobs at once: getting good enough at the work, getting through the interviews, and finding the openings in the first place. The first two are worth the effort. The third mostly isn’t: it’s refreshing hundreds of career pages, scrolling past senior roles and clearance-only postings, and finding the good one a week after it went up.
So I built the thing I wanted to read each morning: the new roles that fit me, ranked, with a one-line reason for each and nothing I’d have to rule out myself. It’s built for speed to apply, because an early application is the one most likely to get read.
One rule shapes every decision in it: a missed role is worse than an extra one. A false positive costs a line in a digest. A false negative loses a real job, and I’d never know it existed.
What lands in the inbox
Every run ranks roles by how recently they were posted and how well they fit my résumé, then writes a digest: which résumé to send, what the posting actually requires, and why it’s a match. On October 2 it recommended 24 roles out of 4,447 that qualified.
24 recommended of 4,447 qualifying
Internships & co-ops
Software Engineering Intern — Natera
High fitRésumé: intern-2028, so it reads as a returning student
Why Scripting, schema changes, data validation and unit tests line up with your Python/SQL work, and Google Sheets Apps Script is a stated plus. It is remote within the USA, so no relocation is needed.
Software Engineer Intern — ShopBack
High fitRésumé: intern-2028
Why The posting rewards things built outside class and applied AI work; your arbitrage platform and a live, tested NYCHappyHours site are that kind of evidence. NYC is a listed location.
Spending money only where it decides something
Reading every posting with a frontier model would work, and it would cost far too much to run every hour for anyone but me. So each role passes through a funnel where every tier is cheaper than the one after it, and a role only moves on when the cheaper tier can’t decide.
- Every role on 740+ career sites, plus community feeds swept as often as hourly
- 25-rule deterministic screen $0 Rules out roles on what they state: seniority, years of experience, clearance, licensure, an advanced degree, the graduation window, location.
- A classifier API reads the posting’s facts about $0.05 a run Only for the facts the rules couldn’t settle from the text they had.
- Claude scores fit and writes the summary top-ranked roles only Through the API, or through an AI subscription you already pay for.
- The morning digest under $0.50 per user per day, projected
The cheapest win wasn’t clever: it was reading the classifier’s API docs properly. Asking all of a posting’s questions in one request instead of one request per question, then skipping duplicate text and repeated boilerplate, cut the projected cost of re-checking 100,000 postings from $32.09 to $2.81. The boilerplate stripping deliberately keeps any sentence that could disqualify a role, so the shortcut can cost a line of noise but never a real job.
Bring your own AI subscription
The last step, the fit review, doesn’t have to run on an API key at all. Plenty of people job hunting already pay for Claude or ChatGPT, so Role Scout can hand that step to the subscription they have, which keeps both cost and setup time about as low as they go: no API keys to create, no billing to manage, no new tool to learn. Each run writes a brief: the ranked roles, the rubric, and your résumé. You give it to a Claude or ChatGPT/Codex session, it writes its judgments to a file, and Role Scout imports them.
The import holds the subscription to the same bar as the API. Every row is checked against the same schema and rubric, every role id is matched to a real posting, and a skipped role or duplicate rank is rejected. The digest that comes out is identical either way; the review just costs nothing extra on top of a subscription you were already paying for.
Through the API
- Run sends the top-ranked roles to Claude
- Billed per token, on your API key
Through your own subscription
- Run writes the brief: roles, rubric, résumé
- Your Claude or ChatGPT session judges it
- No API key, no extra spend
Same rubric · same schema check · same morning digest
Running on a schedule, never silently dropping a role
A fast tier runs every hour, and two full sweeps a day walk every board end to end. A board the run doesn’t reach in time is deferred to the next run, never quietly skipped.
Every stage that can remove a role has to account for it. Each one returns a receipt, and the run refuses to continue if the numbers don’t add up:
took_in = gave_out + Σ droppedby rule + deferred
Drops are keyed by the rule that made them, never a bare count, so “why wasn’t this role in my digest?” always has an answer that names a rule. Behind it, a SQLite store caches model results by a hash of the posting’s content, so an unchanged posting isn’t paid for twice, and 1,750+ tests hold the whole thing in place.
Calibrating the classifier on real postings
Before trusting the classifier with anything, I ran it against 112 labeled postings and kept only the signals that earned it. The best was whether a posting needs a security clearance you already hold: its scores split into two groups with nothing in between.
That gap mattered because the rule was wrong sometimes. It threw out a Leidos junior software engineering role whose posting asked for an active clearance or the ability to obtain one. With the classifier able to overrule the rule, the first production run rescued 27 roles and confirmed 81 as genuine blocks; graded blind afterwards, 26 of the 27 rescues were right.
Then the full corpus humbled the benchmark. The band it had shown empty wasn’t, and two of the three postings sitting just above the cutoff were wrong rejects: Boeing and Northrop roles that also said “ability to obtain.” I raised the threshold the same day. I also cut three signals that carried no information — one varied by 0.003 across every label — and dropped the idea of letting the classifier rank fit at all: one posting labeled a reject scored 2.21, above the best exceptional match at 2.03. Fit stays with Claude.
Try Role Scout
This is an early version that isn’t fully rolled out yet. Leave your information and I’ll send you a totally free digest of matching roles the day of your scan — expect a short wait.
What I learned
Getting the error direction right mattered more than getting the model right. Once I decided that a missed role is the expensive mistake, most design questions answered themselves: blank locations are kept, unreadable postings go to review instead of the bin, and a cheaper tier may only rule a role out on positive evidence in the posting, never on missing information.
The other lesson was the same one the arbitrage project taught me: measure before you trust. The benchmark said a threshold was safe; the full corpus said otherwise. A signal that looked useful carried nothing. Each time, the honest number was the cheaper one to hear early.