Field notes from Jevathon, Jev’s first community hackathon, at CodeRabbit in San Francisco.
On Saturday, September 26, eleven days after TypeSafe AI left stealth with $40M in seed funding and Jev as its first public model, a full room of builders gathered at CodeRabbit to hack on it. Hacking ran from 11:30 AM to 2:30 PM, three hours total. Founding developer advocate Allie Laabs put Jev’s Discord at 110,000 members in its first week, up from roughly zero. TypeSafe AI, she added btw, is hiring.

The name carries a joke with big teeth. TypeSafe.ai named Jev after William Stanley Jevons, the economist who argued in 1865 that more efficient steam engines would raise Britain’s coal consumption. TypeSafe’s bet runs in the same direction for judgment. Make a decision cheap enough and software starts making far more of them.
Decisions Without Sentences
Jev writes no prose. TypeSafe calls it the first System One model, a name drawn from the fast, intuitive System 1 thinking Daniel Kahneman popularized. A request pairs a block of state, whether a string, a JSON object or an array of text, with one or more typed questions. Three primitives shape every answer. Choice selects one option from as many as 255, Score places the state on an ordered scale of 2 to 10 levels and Noul answers a yes or no statement with a single float between 0 and 1. Input costs $42 per billion tokens and output costs like nothing.

Alex Warren of TypeSafe traced the design back to training. Chat models, he explained, learn under a twin goal: pleasing people through RLHF while chasing verifiable rewards in synthetic environments. CEO Diogo Almeida, a former OpenAI researcher, co-invented RLHF and ChatGPT. TypeSafe’s own pitch holds that RLHF-trained models aim to please and assist people, which works against reliable autonomous decisions.
Warren offered practical guidance. Asked whether a 0.7 means 70 percent, he described a rubric: a Score works like a Likert scale where the developer defines each point, and the only reliable calibration check is reading real outputs against the use case. Jev always answers, so expecting a non-answer amounts to a type error. To catch an invalid question, ask a second one about validity or add “not applicable” as a Choice option. Chat models reason one token at a time; with Jev, reasoning lives in a decision tree the developer engineers, or in a hybrid with an LLM. Two documents belong in two requests, since packing both into one payload doubles the questions and splits the model’s attention. Questions drafted by an LLM tend to be overcomplicated, so simpler wording tends to score better.
On cost, Warren went into detail. Billing runs per input token through a proprietary tokenizer, so callers estimate beforehand; every response returns the exact token count, and cost follows from the published rate.
For an agent paying its own way, that means estimate first, reconcile after. I asked him about the workload for my hack for around 720 states and 30 questions. Warren guessed, probably costs under ten cents. And mentioned that batch endpoints may come, aimed at enterprise loads of hundreds of thousands of states.
“(Coding) Agents tend to think it’s more expensive than it is,” Warren said.
Next on the public roadmap: image input, already under internal testing after two years of building in stealth.
Judgment, Not Arithmetic
Laabs opened with a boundary. Jev is not an agentic product, she told one hacker, and does nothing until a programmer puts it into software. Hand an agent Jev with a mandate to decide and results wobble, since agents struggle with consistency. TypeSafe builds harnesses with Jev inside the loop, choosing tool calls and assessing tool-call safety.
Her pattern for repeatable work: collapse the process into a single tool call, mostly deterministic code, with Jev at the few probabilistic joints. Software crawls an asset folder; Jev judges whether a new file updates an old one. The hacker proposed a meta-hook that splits each incoming request into atomic decisions before farming them out to Jev. Laabs told him to explore it and pointed to Jevify, a community skill that scans a codebase for places Jev could replace LLM calls.
Enterprise adoption, she said, skews classification-shaped. Many teams had already proved an LLM could categorize their inputs, then shelved the prototype when production volume erased the ROI. Cheaper decisions revived those projects. Fraud detection and cybersecurity share the shape, down to scoring network packets in real time at finer granularity than before. In narrow domains a fine-tuned classifier may still win, she conceded, though keeping one current costs labor and money.

Her filter for any idea starts with the input. Pure numbers usually point to arithmetic, and arithmetic belongs on a CPU. Her own Doom demo, built on entity positions, would lose to conventional code as a serious bot. Jev becomes the right tool when spoken input needs turning into a decision or an opponent needs a probabilistic personality. Taste works the same way: decompose it, compute what color science already knows and reserve judgment for what remains. She learned the flip side from a quick tarot experiment. Without knowing tarot, she couldn’t tell whether the output was any good.
“If it’s vibes, now we’re getting somewhere,” Laabs said.
Five Minutes, No Slides
Rules stayed strict: five minutes per team, microphone cut at the limit, live working products only and $1,000 for first place. The published panel drew judges from Anthropic, xAI, Photon and CodeRabbit alongside startup founders and early-stage investors. One judge, Lucas Gonzalez Pagliere, a research PM at Anthropic focused on computer use and agentic systems, said a clip of Jev driving a car in GTA 5 convinced him to attend.
A dozen people spent the afternoon judging software built on a machine that judges.

First Place: Jevolution
CodeRabbit interns Utkarsh Gupta, Xirui Huang and Nikhil Hooda built a multi-agent ecosystem simulator for wildlife researchers studying endangered or data-deficient species without experimenting on live animals. Wolves and rabbits decide in real time as temperature, humidity, food supply and danger signals shift, with Jev driving each choice. The team claimed Jev handled about 4,000 decisions for roughly $1.30, against about 500 for roughly $12 with Claude. Per decision, that works out to about three hundredths of a cent against nearly two cents. The comparison measured cost, not decision quality.
Allie Laabs’s may ask the question: how much of a wolf’s hunt reduces to arithmetic, and how much to vibes?

Second Place: Jev Shopping Network
Home shopping TV crossed with the interdimensional cable of Rick and Morty. Dials for weirdness and tackiness steer Jev as it scores arbitrary catalog items in real time, surfacing finds like a sculpture of a dog in meditation pose. Some agentic commerce systems grade product claims for truth. This one graded them for tackiness, deliberately, through a rubric this second-place winner wrote.

Third Place: Pixie Dust
Rick Lopez, a developer with 12+ years of experience building back-office automation at a stealth startup, records browser flows, from fetching sports scores to back-office chores. Jev maps decision points and variants into production-ready API endpoints, reachable by CLI or API. It reads like Laabs’s advice turned into a product: collapse a repeated process into one call, with judgment only where needed.

Final Two Demos — honorable mention
First Minutes missed the podium but carried the day’s highest stakes. Untreated, a stroke kills an estimated 1.9 million neurons a minute. EMS crews lose time to long paper checklists and to working out which hospital can take the patient. First Minutes takes initial vitals, such as age, facial droop, eye deviation, arm weakness and time last seen normal, runs them through a clinical decision tree with Jev and asks only the follow-up questions that matter. Gemini queries hospital bed and surgery availability to pick a destination.
One judge pressed on hallucination: should a model route a stroke patient at all? The team answered that Jev scores a fixed decision tree instead of generating free text, works in seconds where stressed crews spend twenty-plus minutes on forms and checks dozens of conditions in parallel. A typed answer can’t invent a hospital. It can still choose the wrong one, a gap between schema-valid and correct that TypeSafe’s CEO acknowledged in the Hacker News launch thread. Remember what Warren said: Jev always answers. In an ambulance, that matters. Escalation has to be designed in, and TypeSafe attaches a confidence estimate to every decision so software can act when confidence runs high and escalate when it doesn’t.
Judges later credited the team’s technical depth and real-world utility, but noted it showed a static text interface where its live routing map would have carried the pitch. The most consequential demo of the day lost ground on presentation.

GitHub Roast came from Kevin and teammates, first-time hackers from Texas. The tool pulls a developer’s GitHub history, from repos and PRs to forks and commit times, and roasts it at a chosen spice level. Wi-Fi failed on the first attempt and the host sent them offstage. Over two more attempts they roasted a fellow hacker live with voice output, down to his commit hours, follower count and account age. Since Jev returns no text, its role sat at the judgment layer, with the jokes delivered elsewhere in the pipeline.

Host With 1,500 Open PRs
CodeRabbit hosted, and its own story fit the theme. DevEx engineer Hendrik Krack described the company unleashing background agents on its monorepo overnight.
“We woke up a night later with 1,500 pull requests open, and not enough humans in our company to press the button,” Krack said.
CodeRabbit’s answer, Triage, ranks agent-written PRs by risk and readiness for review. Deciding which PR deserves a human is itself a judgment problem. Thanks to Sourabh Mane, a design engineer at CodeRabbit, previously founding designer at Sieve (YC 22) with research roots in HCI and AI, who also sat on the judging panel.
Trust Gap, Ten Dollars at a Time
The AI Collective, a global non-profit, counts more than 200,000 members across 100-plus forums and describes its work as building the human layer for the AI era. Organizer My Luu framed the mission around a single term.
“Our whole mission is to bridge the gap that’s widening between how fast technology is moving and how humans cannot keep up with the pace of change, so we call that the Trust Gap,” Luu said.
Every event, she said, should serve people, lowering barriers so builders feel empowered instead of intimidated. Her skepticism landed on machine payments. She questioned why any enterprise would let an agent manage its finances when she wouldn’t let one manage hers. My answer to her: start small. Give an agent $10 and one task, like finding a useful subscription under a set price. Trust, perhaps, arrives in allowances long before it arrives in accounts.
Priced to Multiply
William Stanley Jevons watched cheaper coal feed more furnaces. Here at the Jevathon, cheap judgment fed wolves, dog sculptures, API endpoints and ambulance routes, thousands of calls for less than a cup of coffee. Twelve humans needed an afternoon to rank a handful of live demos, weighing a routing map they never saw against a roast that nearly failed to load. Jev would have answered instantly. It always answers.
If judgment becomes nearly free, which questions still deserve someone who hesitates?
Thanks to The AI Collective and all event hosts including AJ Green, My Luu, Roan Weigert, Chappy Asel, and sponsors CodeRabbit, Browserbase, GMI and Photon.
X post 🐦
Sources
- Flavio Copes, “A deep dive into Jev, TypeSafe’s System One model” — independent technical explainer covering Jev’s launch, primitives, latency, pricing and fit.
- TypeSafe AI — official site and launch post “Introducing System One Models & Jev,” covering model design, naming, confidence estimates and pricing.
- Jevathon event page, TypeSafe AI × AI Collective — event listing with schedule, judging panel, hosts and AI Collective mission.
