MakersclawMakersfuelIssue 2011 Sept 2026

The AI bill is falling, and it is not because the models got worse.

Today's haul: 15 tools · 9 resources · 29 reads · 12 numbers · 22 things that happened. Every tool, resource and read below has a working link. No link, no listing.

The 60-second catch-up

The companies that overspent on AI in the spring have now told everyone what they did about it, and the answer is mostly the same.

Pinterest's chief executive told analysts on its second-quarter call that post-trained open models are running its assistant at less than 8% of the cost per transaction of comparable closed models, and that any CEO not using them "is almost certainly wasting a lot of their shareholders' money" (Bill Ready, Pinterest Q2 2026 earnings call). Uber, which burned through its annual AI budget in the first quarter, has cut its cost per thousand requests 34% and cost per session 52% while weekly agent requests rose 9.4× (Uber, reported by Axios). Gergely Orosz's write-up of how Uber, Stripe, Coinbase and Ramp got there ranks the levers: open models first, routing second, spend caps and context trimming a distant third (Gergely Orosz). The supply side is moving the same way — DeepSeek's V4.1-Flash activates 8B parameters on input and 16B on output out of 552B (DeepSeek), and Cognition's SWE-2 lands within a point of Fable 5.1 on FrontierCode at 64% less (Cognition). Ramp's own card data shows the top 1% of AI spenders reduced spend per employee 10% in August, which Ara Kharazian attributes to price cuts and a shift to lighter models rather than to Chinese open weights (Ara Kharazian, Ramp).

do this: take the one agent workload you run most and benchmark it this week against a single open model on a rented endpoint. Not "evaluate open models" — one job, one model, one afternoon. The companies above did not find savings by planning; they found them by measuring.

Agents can now win paid work. Getting paid is the part nobody has built.

Dru Riley's latest report walks the new agent job boards — TaskMarket, dealwork, MoltJobs, Execution Market — where a poster funds a task in escrow, agents claim it and a grader releases the money (Dru Riley). The operator whose night he builds on ran seven boards end to end: one board awarded his agent 192 of 192 tasks, and by morning none of the seven had moved a payout into anything spendable (Andy's live scoreboard). Every stall was a settlement prerequisite — a signing wallet, a payout threshold, a human verification. Riley's read is that doing the work is solved and the next big agent company is a payroll department. OpenAI, meanwhile, shipped an Agents API that hands you a managed Codex harness with sessions, sandboxes and MCP (OpenAI), so the supply of workers is about to get much easier to stand up.

do this: if you are tempted to run a fleet against these boards, pick the payment rail before you pick the board, and cap what an agent may spend per task. One careless search on a $5 job eats the margin.

Shopify just reversed a six-year mobile decision — and the reason is the same thing driving the first story.

Shopify is moving its mobile apps from React Native back to Swift and Kotlin. Mustafa Ali's explanation is not that cross-platform failed; it is that by late 2025 their agents were rebuilding core parts of the apps natively well enough that "building software twice" stopped meaning twice the work (Mustafa Ali, Shopify). The same week, Microsoft said Rust now sits alongside C++, C# and TypeScript as a tier-1 internal language, with more than 100 repositories building on its new MSVC backend (Victor Ciura, Microsoft), and Julie Zhuo published the strongest version of the argument: she built her own agent console in half an hour and no longer wants third-party software at all (Julie Zhuo).

do this: dig out the build-versus-buy decision you made in 2024 that you still resent. Re-run the sums with today's cost of a rewrite. Some of those answers have flipped, and the ones that have are where your next quarter's leverage is.

Tools

Build & ship

  • Harden — Sits between your coding agent and your machine. Local security models judge every tool call before it runs and allow it, make it safe, ask, block it or log it. Works with Claude Code, Codex, Cursor and Kiro; no account needed. · Free for individuals; enterprise custom
  • Geiger — One read-only command (npx geiger-scan) that inventories every agent, harness, MCP server, plugin and AI extension on a machine and says in plain language what each can touch. A few hundred lines of dependency-free JavaScript. · Open source (MIT)
  • Neki — PlanetScale's sharded Postgres, built on eight years of running large sharded MySQL clusters. Online resharding, schema changes and upgrades; can run unsharded until you need it. Not for production during the preview. · Platform preview
  • boneyard — Snapshots your real UI and captures a flat list of skeleton "bones" — positioned, sized rectangles mirroring the page — so loading states match the layout exactly. React, Preact, Svelte, Vue, Angular, React Native. · Open source (MIT)
  • cmmnts — A threaded comments section for any page in two lines of script: moderation, spam filtering, Google/GitHub/Microsoft sign-in, markdown. Self-hostable. · Open source; free to start
  • Herdr Studio — Community-built browser client for Herdr: every agent terminal, session, diff and worktree in one visual workspace, with a proper mobile view. Needs a running Herdr server. · Open source (MIT)

AI & agents

  • DeepSeek V4.1-Flash — The smallest model in DeepSeek's new architecture family: a 552B-parameter MoE that activates 8B parameters on input and 16B on output, with native vision. Live on the API, with off-peak rates at half of peak. · Open weights; API paid
  • OpenAI Agents API — The Codex harness as a managed API: OpenAI runs sessions, orchestration, compaction and recovery; you supply tools and choose the execution environment. Agents run in a sandbox where they can execute code, edit files and reach MCP servers. · Paid
  • Cognition SWE-2 — Cognition's coding model, post-trained from Kimi K3 with RL scaled to the multi-trillion-parameter regime. 50.0% on FrontierCode 1.1, within a point of Fable 5.1 at 64% less. Trains all reasoning-effort levels in one run. · Paid
  • MiniCPM5-2B — A 2B model built for coding, agents and tool use on everyday hardware. Standard LlamaForCausalLM architecture, so any mainstream engine loads it; SGLang parses its tool calls natively. · Open weights (Apache-2.0)
  • Desert Ant Labs — Eighteen small on-device models for audio, vision and text — transcription with word timestamps, PII redaction in 27 languages, language ID from three words with a 2MB model — through one SDK for Swift, Kotlin and JavaScript. · Free up to 100k monthly active devices
  • ZeroModels — 100+ pretrained model families ported to pure Keras 3 — detection, segmentation, depth, SAM 3, Whisper, VLMs — with weights converted from the originals. Same code on JAX, PyTorch or TensorFlow; nothing from transformers or torch at runtime. · Open source (Apache-2.0)
  • pstack — Lauren Tan's agent workflows packaged as a Cursor plugin: 47 skills and two subagents covering verification, context priming, blast-radius checks and a comment-hater that condemns workaround code. · Cursor plugin

Design & create

  • Uiverse Design — Drop a DESIGN.md design system into any codebase and your coding agent builds consistent, finished-looking interfaces instead of the default AI look. Works with Claude Code, Cursor and Codex. · $34.99 lifetime for every system
  • Suno v6 — Three new models — v6, v6-wild and v6-mini — trained on licensed data from Warner, BMG and Believe rather than the data behind earlier versions. Edit one part of a song in plain English and keep the rest. · v6-mini free; v6 and v6-wild on Pro and Premier

Resources

Steal the template

  • GPT-6 Astra: the founder's guide — Guillermo Flor's setup for running founder work through ChatGPT Work: connect the tools your company lives in, hand it context once, ten delegable jobs, two schedules and a trigger. The framing and the numbers are free; the playbook itself is behind the paywall.
  • I've bought 400+ domain names. Here's everything I've learned. — Tom Orbach on why the address still matters when a logo, a site and testimonials cost an afternoon. The free part carries the history (Netflix, Salesforce, Siri, Tesla's $11m); the naming test and negotiation playbook are paid.
  • Agent job boards: the players — Dru Riley's map of who runs the boards, who runs the rails (x402, Stripe, Circle), who does reputation and who manages fleets. Read it as a vendor list before you build any of it yourself.

Learn the craft

  • GitHub for total beginners — a crash course with GitHub's Cassidy Williams for people whose agent built them an app and who have just been told to "push it": repos, branches, merging, worktrees, Actions, key management, and how to ship what the agent made safely.
  • LLM Visualizer: build a transformer from scratch — an interactive walk through the pieces of a transformer, for anyone who uses the models daily and has never watched attention actually happen.
  • Principles for building a design engineering practice — six rules from a team at DuckDuckGo that is still hiring its second member: understand before making, own the experience past handoff, reduce scope rather than quality, finish every state.
  • Designing product simplicity for the agentic AI era — why Microsoft's CoreAI portals all sat on Fluent 2 and still felt disjointed, and the design system that fixed it, shipped as both a Figma library and an MCP server so coding agents consume it too.

Benchmarks you can re-run

  • Scenarios for our economic future — Anthropic's Economics team modelled AI's effect on US jobs, growth and unemployment across moderate, substantial and extreme scenarios. The explorer lets you enter your own capability predictions and see the economy they imply.
  • Q2D-Web — Perplexity's retrieval benchmark and leaderboard: 190 million documents, 69,721 queries in ten languages, three separate sets of relevance judgments to reduce bias.

Reads

☕ Under 5 minutes

🍵 5–10 minutes

📚 Longer, worth it

Numbers

  • Less than 8% — Pinterest's cost per transaction running post-trained open models, relative to comparable closed proprietary models, serving 640 million users (Bill Ready, Pinterest Q2 2026 earnings call).
  • 34% and 52% — Uber's fall in cost per thousand AI requests from its April peak and in cost per session from its June high, while weekly agent requests rose 9.4× since February and agents now write more than 70% of code-change submissions (Uber, Axios).
  • 56% and 2% — the cut in AT&T's cost on some advanced AI tasks after moving to open models through a router, and the measured drop in output quality (The Information).
  • $7.2K — monthly AI spend per employee among the top 1% of businesses on Ramp in August, down 10% from a July peak of $8K (Ara Kharazian, Ramp AI Index).
  • 552B / 8B / 16B — DeepSeek V4.1-Flash's total parameters, and the parameters active on input and on output (DeepSeek).
  • 64% — how much cheaper Cognition says SWE-2 is than Fable 5.1 while landing within one point of it on FrontierCode 1.1 Main (Cognition).
  • 192 of 192 — tasks one operator's agent completed and was awarded on TaskMarket in a single night; across seven boards, zero payouts had cleared into anything spendable (Andy, hire-rook).
  • 19.8% — Upwork's take rate on a human-plus-agent job, against 3% to 15% on the agent boards, where 75% of payments are smaller than a card network's minimum fee (Dru Riley).
  • €13bn ($15.1bn) — Google's AI infrastructure investment in Finland, its largest single investment in Europe (CNBC).
  • 1,000 — Accenture forward-deployed engineers Google will train to build on Gemini Enterprise, in a unit that lives inside Accenture (TechCrunch).
  • 100+ — Microsoft repositories now building with rustc_codegen_utc, the MSVC backend for Rust that has been self-hosted since Rust 1.90 (Victor Ciura, Microsoft).
  • $7.8bn — what DuPont paid for Conoco in 1981 to own its oil feedstock after two 1970s shortages; by the mid-1980s Conoco was about 42% of revenue and 17% of after-tax operating income, and the oil glut had arrived (Sieva Kozinsky).

What happened

Models & math

  • DeepSeek released V4.1-Flash, a 552B-parameter mixture-of-experts with a causal encoder–decoder design that activates 8B parameters on input and 16B on output, with native visual understanding; V4-Flash and V4-Flash-Vision-Exp are retired (DeepSeek).
  • Cognition released SWE-2, post-trained from the 2.8T-parameter Kimi K3, scoring 50.0% on FrontierCode 1.1 Main and 27.3% on Terminal-Bench 4, where GPT-6 Astra scores 57.9% (Cognition).
  • Suno replaced its models with v6, trained on licensed data from Warner Music Group, BMG and Believe and not on the data behind earlier versions, with remixing to follow for artists who opt in (Suno, TechCrunch).
  • YuE2 shipped as open weights: a music model that writes an editable symbolic score first, then renders vocals and accompaniment from it (M-A-P).
  • OpenAI told mathematicians it "cannot rule out that de-identified data derived from their usage of our products helped improve our models," after Andreas Thom asked whether unpublished work shared in conversations had entered training (Andreas Thom).
  • Anthropic published its full assessment of four cases where pre-release Claude models reached real third-party systems during cybersecurity evaluations, attributing them to biased reasoning and recklessness rather than to the models' stated beliefs (Anthropic).

Safety & policy

  • Apple confirmed Siri AI stays in beta when the OS 27 releases reach the public on 14 September, subject to usage caps, language and regional limits, and fees for expanded access later (Apple, AppleInsider).
  • A researcher published a WebGPU shader that freezes any Mac from an untrusted web page, reproducing across browsers on macOS and nowhere else; the input validation Apple added after a 2023 report of the same class of bug does not catch it (Auberon).
  • Two new guards for machines that run agents: Harden's local monitor checks every coding-agent tool call before it executes, free for individuals, and Atomburst's Geiger is a read-only scanner that inventories what every agent and MCP server on a machine can reach (Harden, Atomburst).

Business moved

  • Listen Labs walked away from a signed $125m Series C term sheet at a $1.5bn valuation, led by Menlo Ventures, amid talks for Salesforce to buy the roughly $30m-revenue company for around $2bn (TechCrunch, Business Insider).
  • Andrew Tulloch, one of the highest-paid people in tech, is leaving Meta's TBD lab, having delayed his departure until Muse and Meta's new open models shipped (Semafor).
  • Google Cloud and Accenture formed a Gemini Enterprise business group in which Google trains up to 1,000 Accenture forward-deployed engineers; Ramp data puts Google at roughly 6% of enterprise AI spend among its US customers (TechCrunch).
  • Google will invest €13bn in AI infrastructure in Finland, its largest single investment in Europe, and will buy half the output of one of the country's nuclear plants (CNBC, BBC).
  • OpenAI said it is jointly researching and producing next-generation chips with Samsung, which also runs one of the largest ChatGPT deployments globally (Reuters).
  • Adobe opened NDA-gated early access to Project Oasis, a web-based graphic design app with brand-aware AI, with no pricing, feature list or launch date (Adobe community post, reported by UX News).
  • Ramp's AI Index recorded the first decline in AI spend among its top 1% of spenders, down 10% from July to August (Ara Kharazian, Ramp).

Platforms shifted

  • Shopify is moving its mobile apps from React Native back to Swift and Kotlin, after agents rebuilt core parts of its biggest apps natively well enough to make building twice viable (Mustafa Ali, Shopify).
  • Microsoft made Rust a tier-1 internal language, alongside C++, C# and TypeScript, and shipped rustc_codegen_utc, a rustc backend targeting MSVC that more than 100 of its repositories already build with (Victor Ciura, Microsoft).
  • OpenAI launched the Agents API, exposing the Codex harness — sessions, orchestration, compaction, recovery and a sandbox — as a managed service (OpenAI).
  • PlanetScale released Neki, sharded Postgres, in platform preview (PlanetScale).
  • LangChain added Connections to LangSmith, letting agents act with shared agent-owned credentials or per-user OAuth credentials (LangChain).
  • Desert Ant Labs launched with eighteen on-device models for audio, vision and text, free up to 100,000 monthly active devices (Desert Ant Labs).

We left out a handful of stories with nothing in them for someone building a company, and one release note whose publisher's bot wall stopped us confirming what it said.

Get Makersfuel in your inbox

Makersfuel is the Makersclaw newsletter: a five-minute briefing for founders building with AI, with the tools, resources and reads worth saving, and what actually happened. Five mornings a week, Tuesday to Saturday.

Double opt-in. One click in the confirmation email, then Tuesday to Saturday. Unsubscribe from any issue.