The frontier of innovation, one Friday in San Francisco
Today I spent the day at the Amazon Web Services (AWS) Builder Loft in San Francisco for the Loop Engineering Hackathon — with a deceptively simple challenge: ship a self-directing agent in a day. Not a chatbot with tools. An agent that plans, acts, observes, and self-corrects across a full build cycle.
Sixty-five teams submitted. I’ve now reviewed them all.
I go to a lot of these events, and I’ve learned to watch for the moment when a room full of builders independently converges on the same idea without coordinating. That’s usually the signal that something has moved from research to practice. It happened here, and it wasn’t the model, the framework, or the demo polish.
It was the🚪 gate.

Three trends that resound 🔔
1. Autonomy is being engineered around denial, not capability. The most sophisticated projects didn’t ask “what can my agent do?” They asked “what happens when my agent is told no?” The best demo of the day was an SRE agent whose restart request was rejected by a zero-trust proxy with a real HTTP 403 — and it pivoted to a rollback and recovered production in 24 seconds. Denied actions weren’t failures; they were information that reshaped the plan. Nearly half the field put an identity-aware policy gate (mostly Pomerium) between the agent and anything irreversible.
2. Agents now have wallets — and budgets. The x402 micropayment protocol (USDC on Base, via Zero.xyz’s capability marketplace) showed up everywhere. Agents that discover a capability gap, shop a machine-to-machine marketplace, pay $0.03 per call, and keep receipts. One team gave their agent $5 and told it to buy the Salesforce Tower. It came back with a $2 Amazon gift card and a live storefront — every transaction settled and auditable. Capability-on-demand plus spend policy is quietly becoming the answer to “how do agents get useful without getting dangerous?”
3. Verification is becoming the product. The self-referential cluster was the most technically serious: agents that supervise other agents, harden eval suites, heal training data, and — my favorite — catch a frontier model fabricating data provenance. The strongest teams refused to let the model grade itself. Hard signals only: exit codes, file hashes, settled transactions, re-run benchmarks.
My prediction
Within 18 months, “agent” won’t mean a reasoning loop with tools. It will mean a governed loop: a planner, an executor, an independent verifier, and a policy gate with the authority to say no — with real money and real credentials scoped per action, not per agent. The teams below are already building it.
My five picks (of 65)
Chosen for genuine closed-loop self-correction and verifiable results — not demo polish.
1. Enrichment Diet — the agent that learned to stop paying
Aristarkh Manchuliantsev · Devpost · GitHub

— dashboard: $0.141/candidate, −67%, quality 93/100, live drop-one-retry log
A recruiting agent that buys candidate data from six paid APIs, then runs a drop-one-and-retry optimization to discover the minimal set that still clears a quality bar. Result: cost per candidate from $0.43 → $0.14 (−67%) with quality held at 93/100. Real payments, settled once; the retry loops run on cheap Akash inference.
“Real-dollar feedback changes how you design agents — savings aren’t estimated tokens, they’re settled transactions with hashes.”
“The load-bearing set differs per candidate: for one founder the $0.25 funding API is pure waste.”
Tech: Akash (Llama 3.3-70B) · Pomerium · Zero.xyz x402 · Node/Express
2. melkat — $5 and no script
Katerina Tchilinguirov, Melany Macías · Devpost · GitHub · Demo video (Loom, 50s)

LADDER dashboard: net worth, agent thoughts log, x402 ledger, 7-rung climb
The team funded a Claude agent with $5 of real USDC and pointed it at a machine-to-machine marketplace with one absurd goal: buy the Salesforce Tower. With five tools (search, inspect, buy, review, narrate) and hard spend limits, the agent scouted the market, bought inputs, built a product, opened a shop — and cashed out a real $2 Amazon gift card, machine-bought from machine. Every run diverges because every decision is made live.

“CASHED OUT”: the real $2 Amazon gift card
“We hand an AI agent $5 of real money and it goes shopping in a marketplace built for machines.”
“The brain is Claude making every decision live. There is no script.”
Tech: Claude Agent SDK · Overclock Labs, creators of Akash Network · Nexla · Zero.xyz · x402 · USDC
3. FleetYield — reliability as a scheduling primitive
Manas Sharma, Sanskriti Uma · Devpost · GitHub

live 24-GPU floor - learning loop
Middleware that sits in front of an unmodified GPU scheduler and turns failure risk into placement decisions — draining degrading GPUs, matching “yellow” hardware to checkpointable jobs, and rehabilitating it via canary promotion. Measured on a live Kubernetes cluster running NVIDIA’s KAI scheduler: +45.8% net goodput, failures 3→0, premium-GPU misuse 32→0, with zero scheduler changes.
“Reliability is a scheduling primitive. Treating a GPU’s failure risk as a first-class price reclaims the entire middle of the fleet.”
Tech: Python · FastAPI · Kubernetes · NVIDIA KAI Scheduler · DCGM · Nexla/Pomerium/Zero
4. AEGIS — the 3 a.m. engineer that accepts “no”
Wally Ahmed · Devpost · GitHub · Live deck

“Who watches production at 3 a.m.?” loop diagram → “THE GATE HOLDS: HTTP 403”
An autonomous on-call SRE that watches live telemetry, diagnoses incidents, and repairs them — but never holds root. Its first fix attempt (restart) was denied by a real Pomerium policy with a real 403. It didn’t stall; it proposed a rollback, got it authorized, and recovered the service in 24 seconds with zero human actions.
“Every action must pass through a zero-trust policy gate it cannot bypass.”
“This denial is a real 403 from a real proxy — authority is enforced, not simulated.”
“Recovery in 24 seconds, human actions: zero.”
Tech: Nexla (telemetry) · Pomerium (zero-trust gate) · Zero.xyz (paid incident report, stablecoin on Base)

5. StageLoop — the agent that critiques its own work (and pays its own bills)
Ashok Sravanam · Devpost · GitHub

3D apartment - run history with scored iterations (3/10 → 8/10)
Give it a real listing URL and $5. It scrapes the apartment (paying per-scrape via x402 from its own wallet), rebuilds the floor plan in 3D, stages your furniture at true scale, then critiques its own placement and loops until it scores ≥7/10 with zero geometric faults. The improvement arc on a real Craigslist listing — 3/10 → 3/10 → 8/10 accepted — is fully auditable in the run history.
“Moving is a leap of faith: you sign a lease without knowing if your bed even fits.”
“Loop until score ≥ 7 with zero geometric faults (max 3 variants).”
Tech: Akash · Blender · FastAPI · Three.js · Nexla · x402/Zero

The close
A year ago, “agent” meant a chatbot that could call an API. Watching 65 teams build in a single day at the Builder Loft, the definition has moved: an agent is now a loop with a conscience — it plans, it acts, it checks its own work against the world (not its own opinion of its work), and, crucially, it accepts being told no.
The teams that internalized that last part didn’t just demo well. They built things you could imagine running unattended — in production, with a wallet, at 3 a.m.
That’s the frontier. It’s not louder models. It’s better loops.
Thanks to judges & speakers:
Michael Ludden Zero.xyz · Abhijit Bharadwaj & Amey Desai Nexla · Nick Taylor Pomerium · Greg Osuri Overclock Labs, creators of Akash Network · Nicholas M. Metaview · Keir L. Filmore · Paul Klitzke Strala · Marmik Patel Meta · George Gulabyan Crusoe · Siddhant Poojary Google · Carlos Rufo Alchymos · Oscar Brisset Remy AI (YC W26)
And gracious hosts! Marlene Ronstedt (emcee), Fatima Guadalupe Lopez, Jacopo P., Alessandro A. · tokens&