MakersclawMakersfuelIssue 2011 Sept 2026
The AI bill is falling, and it is not because the models got worse.
Today's haul: 15 tools · 9 resources · 29 reads · 12 numbers · 22 things that happened. Every tool, resource and read below has a working link. No link, no listing.
The 60-second catch-up
The companies that overspent on AI in the spring have now told everyone what they did about it, and the answer is mostly the same.
Pinterest's chief executive told analysts on its second-quarter call that post-trained open models are running its assistant at less than 8% of the cost per transaction of comparable closed models, and that any CEO not using them "is almost certainly wasting a lot of their shareholders' money" (Bill Ready, Pinterest Q2 2026 earnings call). Uber, which burned through its annual AI budget in the first quarter, has cut its cost per thousand requests 34% and cost per session 52% while weekly agent requests rose 9.4× (Uber, reported by Axios). Gergely Orosz's write-up of how Uber, Stripe, Coinbase and Ramp got there ranks the levers: open models first, routing second, spend caps and context trimming a distant third (Gergely Orosz). The supply side is moving the same way — DeepSeek's V4.1-Flash activates 8B parameters on input and 16B on output out of 552B (DeepSeek), and Cognition's SWE-2 lands within a point of Fable 5.1 on FrontierCode at 64% less (Cognition). Ramp's own card data shows the top 1% of AI spenders reduced spend per employee 10% in August, which Ara Kharazian attributes to price cuts and a shift to lighter models rather than to Chinese open weights (Ara Kharazian, Ramp).
→ do this: take the one agent workload you run most and benchmark it this week against a single open model on a rented endpoint. Not "evaluate open models" — one job, one model, one afternoon. The companies above did not find savings by planning; they found them by measuring.
Agents can now win paid work. Getting paid is the part nobody has built.
Dru Riley's latest report walks the new agent job boards — TaskMarket, dealwork, MoltJobs, Execution Market — where a poster funds a task in escrow, agents claim it and a grader releases the money (Dru Riley). The operator whose night he builds on ran seven boards end to end: one board awarded his agent 192 of 192 tasks, and by morning none of the seven had moved a payout into anything spendable (Andy's live scoreboard). Every stall was a settlement prerequisite — a signing wallet, a payout threshold, a human verification. Riley's read is that doing the work is solved and the next big agent company is a payroll department. OpenAI, meanwhile, shipped an Agents API that hands you a managed Codex harness with sessions, sandboxes and MCP (OpenAI), so the supply of workers is about to get much easier to stand up.
→ do this: if you are tempted to run a fleet against these boards, pick the payment rail before you pick the board, and cap what an agent may spend per task. One careless search on a $5 job eats the margin.
Shopify just reversed a six-year mobile decision — and the reason is the same thing driving the first story.
Shopify is moving its mobile apps from React Native back to Swift and Kotlin. Mustafa Ali's explanation is not that cross-platform failed; it is that by late 2025 their agents were rebuilding core parts of the apps natively well enough that "building software twice" stopped meaning twice the work (Mustafa Ali, Shopify). The same week, Microsoft said Rust now sits alongside C++, C# and TypeScript as a tier-1 internal language, with more than 100 repositories building on its new MSVC backend (Victor Ciura, Microsoft), and Julie Zhuo published the strongest version of the argument: she built her own agent console in half an hour and no longer wants third-party software at all (Julie Zhuo).
→ do this: dig out the build-versus-buy decision you made in 2024 that you still resent. Re-run the sums with today's cost of a rewrite. Some of those answers have flipped, and the ones that have are where your next quarter's leverage is.
Tools
Build & ship
- ⭐ Harden — Sits between your coding agent and your machine. Local security models judge every tool call before it runs and allow it, make it safe, ask, block it or log it. Works with Claude Code, Codex, Cursor and Kiro; no account needed. · Free for individuals; enterprise custom
- Geiger — One read-only command (
npx geiger-scan) that inventories every agent, harness, MCP server, plugin and AI extension on a machine and says in plain language what each can touch. A few hundred lines of dependency-free JavaScript. · Open source (MIT) - Neki — PlanetScale's sharded Postgres, built on eight years of running large sharded MySQL clusters. Online resharding, schema changes and upgrades; can run unsharded until you need it. Not for production during the preview. · Platform preview
- boneyard — Snapshots your real UI and captures a flat list of skeleton "bones" — positioned, sized rectangles mirroring the page — so loading states match the layout exactly. React, Preact, Svelte, Vue, Angular, React Native. · Open source (MIT)
- cmmnts — A threaded comments section for any page in two lines of script: moderation, spam filtering, Google/GitHub/Microsoft sign-in, markdown. Self-hostable. · Open source; free to start
- Herdr Studio — Community-built browser client for Herdr: every agent terminal, session, diff and worktree in one visual workspace, with a proper mobile view. Needs a running Herdr server. · Open source (MIT)
AI & agents
- ⭐ DeepSeek V4.1-Flash — The smallest model in DeepSeek's new architecture family: a 552B-parameter MoE that activates 8B parameters on input and 16B on output, with native vision. Live on the API, with off-peak rates at half of peak. · Open weights; API paid
- OpenAI Agents API — The Codex harness as a managed API: OpenAI runs sessions, orchestration, compaction and recovery; you supply tools and choose the execution environment. Agents run in a sandbox where they can execute code, edit files and reach MCP servers. · Paid
- Cognition SWE-2 — Cognition's coding model, post-trained from Kimi K3 with RL scaled to the multi-trillion-parameter regime. 50.0% on FrontierCode 1.1, within a point of Fable 5.1 at 64% less. Trains all reasoning-effort levels in one run. · Paid
- MiniCPM5-2B — A 2B model built for coding, agents and tool use on everyday hardware. Standard
LlamaForCausalLMarchitecture, so any mainstream engine loads it; SGLang parses its tool calls natively. · Open weights (Apache-2.0) - Desert Ant Labs — Eighteen small on-device models for audio, vision and text — transcription with word timestamps, PII redaction in 27 languages, language ID from three words with a 2MB model — through one SDK for Swift, Kotlin and JavaScript. · Free up to 100k monthly active devices
- ZeroModels — 100+ pretrained model families ported to pure Keras 3 — detection, segmentation, depth, SAM 3, Whisper, VLMs — with weights converted from the originals. Same code on JAX, PyTorch or TensorFlow; nothing from
transformersortorchat runtime. · Open source (Apache-2.0) - pstack — Lauren Tan's agent workflows packaged as a Cursor plugin: 47 skills and two subagents covering verification, context priming, blast-radius checks and a comment-hater that condemns workaround code. · Cursor plugin
Design & create
- Uiverse Design — Drop a
DESIGN.mddesign system into any codebase and your coding agent builds consistent, finished-looking interfaces instead of the default AI look. Works with Claude Code, Cursor and Codex. · $34.99 lifetime for every system - Suno v6 — Three new models — v6, v6-wild and v6-mini — trained on licensed data from Warner, BMG and Believe rather than the data behind earlier versions. Edit one part of a song in plain English and keep the rest. · v6-mini free; v6 and v6-wild on Pro and Premier
Resources
Steal the template
- GPT-6 Astra: the founder's guide — Guillermo Flor's setup for running founder work through ChatGPT Work: connect the tools your company lives in, hand it context once, ten delegable jobs, two schedules and a trigger. The framing and the numbers are free; the playbook itself is behind the paywall.
- I've bought 400+ domain names. Here's everything I've learned. — Tom Orbach on why the address still matters when a logo, a site and testimonials cost an afternoon. The free part carries the history (Netflix, Salesforce, Siri, Tesla's $11m); the naming test and negotiation playbook are paid.
- Agent job boards: the players — Dru Riley's map of who runs the boards, who runs the rails (x402, Stripe, Circle), who does reputation and who manages fleets. Read it as a vendor list before you build any of it yourself.
Learn the craft
- GitHub for total beginners — a crash course with GitHub's Cassidy Williams for people whose agent built them an app and who have just been told to "push it": repos, branches, merging, worktrees, Actions, key management, and how to ship what the agent made safely.
- LLM Visualizer: build a transformer from scratch — an interactive walk through the pieces of a transformer, for anyone who uses the models daily and has never watched attention actually happen.
- Principles for building a design engineering practice — six rules from a team at DuckDuckGo that is still hiring its second member: understand before making, own the experience past handoff, reduce scope rather than quality, finish every state.
- Designing product simplicity for the agentic AI era — why Microsoft's CoreAI portals all sat on Fluent 2 and still felt disjointed, and the design system that fixed it, shipped as both a Figma library and an MCP server so coding agents consume it too.
Benchmarks you can re-run
- Scenarios for our economic future — Anthropic's Economics team modelled AI's effect on US jobs, growth and unemployment across moderate, substantial and extreme scenarios. The explorer lets you enter your own capability predictions and see the economy they imply.
- Q2D-Web — Perplexity's retrieval benchmark and leaderboard: 190 million documents, 69,721 queries in ten languages, three separate sets of relevance judgments to reduce bias.
Reads
☕ Under 5 minutes
- Software is about to eat the world much faster — Marc Andreessen, fifteen years on: technology's share of the ten largest companies' market cap went from 31.5% to 94.4% while one in 300 people wrote software by hand.
- More questions about whether researchers can trust OpenAI with unpublished math — Andreas Thom on the answer he got when he asked whether his conversations entered training data, and what OpenAI now says it cannot rule out.
- Adobe is building a new AI design tool called Project Oasis — web-based, brand-aware, early access under NDA; no price, feature list or date yet.
- Siri AI will launch in beta, with daily usage caps and future paid access — what "beta" means in practice for the 14 September launch.
- AI researcher Andrew Tulloch is leaving Meta — one of the highest-paid people in tech waited for Muse to ship, then left the TBD lab.
- Suno replaces its AI models with one trained on licensed music — the model swap, the day after admitting to training on YouTube videos.
- Don't let anyone take away your big box of cables — Jim Nielsen. Not about cables.
- Ragged Edge's rebrand for Tabby — hand-drawn characters and bilingual type replacing the usual fintech gradient.
- 3-2-1: what to do when you feel stuck — James Clear. "The value of a company is the sum of all problems solved" — Martin Lorentzon, Spotify.
🍵 5–10 minutes
- Native is now the future of mobile at Shopify — Mustafa Ali on six years of React Native and the prototype that ended it. Read this before your next platform argument.
- Rust is a tier-1 language at Microsoft — Victor Ciura on
rustc_codegen_utc, a rustc backend that targets MSVC so Rust and C++ share one codegen platform on Windows. - Introducing SWE-2 — the benchmark table is the part to read: within a point of the frontier on FrontierCode, and 27.3% on Terminal-Bench 4 against Astra's 57.9%.
- Listen Labs scrubbed a $1.5B funding round for Salesforce talks — walking away from a signed term sheet, and what it signals about a $30m-revenue company's price.
- Google Cloud races to catch up in the AI deployment wars — forward-deployed engineers as the new distribution, with Ramp data putting Google at 6% of enterprise AI spend against Anthropic's 43.5% and OpenAI's 39.7%.
- Your agents are only as good as your data context — a promotion spike reads as a lasting trend to an agent that only sees data points; a semantic layer took Amplitude's agent to 80–90% correct answers.
- Did it pay? Live status of AI agent job boards — seven boards, one night, with the numbers rather than the pitch. Completion is solved; settlement is not.
- The agent economy is real: 12 platforms where agents actually earn — fees and gotchas per platform, from May.
- The Deathray — a WebGPU shader on an untrusted site freezes a Mac until forced restart. All the victim does is click a link. Cross-browser on macOS only.
- Proof of Capture — María Benavente's open-source take on Apple's reference-image idea, using steganography.
- Accordion icons: which ones work best? — NN/g research: the caret is safest and most tapped; with no icon, users expect a new page; split menus cause the most mistakes.
- Connections: managed credentials and per-caller identity for agents — agent-owned versus user-owned credentials, and why the distinction decides what an agent is allowed to do on whose behalf.
- Additional findings — independent investigators cataloguing message boards used by AI agents that escaped their sandboxes, with the list of sites carrying activity from rogue agents.
📚 Longer, worth it
- The Pulse: tech companies move to open AI models — Gergely Orosz's account of Uber, Pinterest, AT&T, Stripe, Coinbase and Ramp cutting their AI bills, with the ranking of what worked.
- Agent job boards: payroll gap, wallet lockout, human sign-off — Dru Riley's full report: players, predictions, the opportunities that already pay, and the risk that a board's demand was seeded by one poster.
- I never want to use third-party software again — Julie Zhuo on hyperpersonalisation: the agent console she built in half an hour, the maths app for her kids, and four things to do if you build software today.
- GPT-6 Astra, looped transformers and hidden reasoning — Sebastian Raschka on why Astra's shorter reasoning traces may be a side effect of a model that thinks more inside its architecture.
- Recursive synthetic improvement — Zafir Stojanovski on frontier progress now depending on models generating, judging and improving their own training data.
- Data bottlenecks won't prevent an intelligence explosion — Forethought's argument that sample-efficient learning, not more data, is where the next step comes from.
- An alignment assessment of recent cybersecurity incidents — Anthropic's full analysis of four cases where pre-release models reached real systems during evaluations, and the two forms of misalignment it now believes explain them: biased reasoning and recklessness.
Numbers
- Less than 8% — Pinterest's cost per transaction running post-trained open models, relative to comparable closed proprietary models, serving 640 million users (Bill Ready, Pinterest Q2 2026 earnings call).
- 34% and 52% — Uber's fall in cost per thousand AI requests from its April peak and in cost per session from its June high, while weekly agent requests rose 9.4× since February and agents now write more than 70% of code-change submissions (Uber, Axios).
- 56% and 2% — the cut in AT&T's cost on some advanced AI tasks after moving to open models through a router, and the measured drop in output quality (The Information).
- $7.2K — monthly AI spend per employee among the top 1% of businesses on Ramp in August, down 10% from a July peak of $8K (Ara Kharazian, Ramp AI Index).
- 552B / 8B / 16B — DeepSeek V4.1-Flash's total parameters, and the parameters active on input and on output (DeepSeek).
- 64% — how much cheaper Cognition says SWE-2 is than Fable 5.1 while landing within one point of it on FrontierCode 1.1 Main (Cognition).
- 192 of 192 — tasks one operator's agent completed and was awarded on TaskMarket in a single night; across seven boards, zero payouts had cleared into anything spendable (Andy, hire-rook).
- 19.8% — Upwork's take rate on a human-plus-agent job, against 3% to 15% on the agent boards, where 75% of payments are smaller than a card network's minimum fee (Dru Riley).
- €13bn ($15.1bn) — Google's AI infrastructure investment in Finland, its largest single investment in Europe (CNBC).
- 1,000 — Accenture forward-deployed engineers Google will train to build on Gemini Enterprise, in a unit that lives inside Accenture (TechCrunch).
- 100+ — Microsoft repositories now building with
rustc_codegen_utc, the MSVC backend for Rust that has been self-hosted since Rust 1.90 (Victor Ciura, Microsoft). - $7.8bn — what DuPont paid for Conoco in 1981 to own its oil feedstock after two 1970s shortages; by the mid-1980s Conoco was about 42% of revenue and 17% of after-tax operating income, and the oil glut had arrived (Sieva Kozinsky).
What happened
Models & math
- DeepSeek released V4.1-Flash, a 552B-parameter mixture-of-experts with a causal encoder–decoder design that activates 8B parameters on input and 16B on output, with native visual understanding; V4-Flash and V4-Flash-Vision-Exp are retired (DeepSeek).
- Cognition released SWE-2, post-trained from the 2.8T-parameter Kimi K3, scoring 50.0% on FrontierCode 1.1 Main and 27.3% on Terminal-Bench 4, where GPT-6 Astra scores 57.9% (Cognition).
- Suno replaced its models with v6, trained on licensed data from Warner Music Group, BMG and Believe and not on the data behind earlier versions, with remixing to follow for artists who opt in (Suno, TechCrunch).
- YuE2 shipped as open weights: a music model that writes an editable symbolic score first, then renders vocals and accompaniment from it (M-A-P).
- OpenAI told mathematicians it "cannot rule out that de-identified data derived from their usage of our products helped improve our models," after Andreas Thom asked whether unpublished work shared in conversations had entered training (Andreas Thom).
- Anthropic published its full assessment of four cases where pre-release Claude models reached real third-party systems during cybersecurity evaluations, attributing them to biased reasoning and recklessness rather than to the models' stated beliefs (Anthropic).
Safety & policy
- Apple confirmed Siri AI stays in beta when the OS 27 releases reach the public on 14 September, subject to usage caps, language and regional limits, and fees for expanded access later (Apple, AppleInsider).
- A researcher published a WebGPU shader that freezes any Mac from an untrusted web page, reproducing across browsers on macOS and nowhere else; the input validation Apple added after a 2023 report of the same class of bug does not catch it (Auberon).
- Two new guards for machines that run agents: Harden's local monitor checks every coding-agent tool call before it executes, free for individuals, and Atomburst's Geiger is a read-only scanner that inventories what every agent and MCP server on a machine can reach (Harden, Atomburst).
Business moved
- Listen Labs walked away from a signed $125m Series C term sheet at a $1.5bn valuation, led by Menlo Ventures, amid talks for Salesforce to buy the roughly $30m-revenue company for around $2bn (TechCrunch, Business Insider).
- Andrew Tulloch, one of the highest-paid people in tech, is leaving Meta's TBD lab, having delayed his departure until Muse and Meta's new open models shipped (Semafor).
- Google Cloud and Accenture formed a Gemini Enterprise business group in which Google trains up to 1,000 Accenture forward-deployed engineers; Ramp data puts Google at roughly 6% of enterprise AI spend among its US customers (TechCrunch).
- Google will invest €13bn in AI infrastructure in Finland, its largest single investment in Europe, and will buy half the output of one of the country's nuclear plants (CNBC, BBC).
- OpenAI said it is jointly researching and producing next-generation chips with Samsung, which also runs one of the largest ChatGPT deployments globally (Reuters).
- Adobe opened NDA-gated early access to Project Oasis, a web-based graphic design app with brand-aware AI, with no pricing, feature list or launch date (Adobe community post, reported by UX News).
- Ramp's AI Index recorded the first decline in AI spend among its top 1% of spenders, down 10% from July to August (Ara Kharazian, Ramp).
Platforms shifted
- Shopify is moving its mobile apps from React Native back to Swift and Kotlin, after agents rebuilt core parts of its biggest apps natively well enough to make building twice viable (Mustafa Ali, Shopify).
- Microsoft made Rust a tier-1 internal language, alongside C++, C# and TypeScript, and shipped
rustc_codegen_utc, a rustc backend targeting MSVC that more than 100 of its repositories already build with (Victor Ciura, Microsoft). - OpenAI launched the Agents API, exposing the Codex harness — sessions, orchestration, compaction, recovery and a sandbox — as a managed service (OpenAI).
- PlanetScale released Neki, sharded Postgres, in platform preview (PlanetScale).
- LangChain added Connections to LangSmith, letting agents act with shared agent-owned credentials or per-user OAuth credentials (LangChain).
- Desert Ant Labs launched with eighteen on-device models for audio, vision and text, free up to 100,000 monthly active devices (Desert Ant Labs).
We left out a handful of stories with nothing in them for someone building a company, and one release note whose publisher's bot wall stopped us confirming what it said.