AI Engineer Miami (April 20, 2026) was a full-day single-track conference gathering engineers from 23 countries. G2i founder Gabe Greenberg opened by announcing Orchestrator AI (orc.ai), a multi-agent orchestration platform.
Benchmarked on SWE-bench Pro against GPT-5.4, it delivered 17.1% lift on easy tasks, 14.
8% on medium, and 8.4% overall — exceeding Opus 4.7 with the same base model.
Dax Raad (OpenCode/Anomaly) keynoted on product restraint: removing engineering bottlenecks strips away the natural filtering of bad ideas, causing rapid product bloat and rot. He urged engineers to absorb multiple user problems before shipping, targeting one high-leverage solution rather than reacting feature-by-feature, and recommended a cadence of one genuinely good idea every few months regardless of AI velocity. Dexter Horthy (HumanLayer) revisited the RPI (Research, Plan, Implement) methodology after deploying it to thousands of engineers.
Three lessons: (1) giving the model the task description rather than neutral questions steers research toward the model's preferred solution; (2) monolithic prompts with 85+ instructions exceed the instruction budget — frontier models reliably follow ~150–200 instructions before adherence degrades, so splitting into sub-prompts under 40 instructions each is essential; (3) reading AI-generated plans instead of code is net-negative work. Shashank Goyal (OpenRouter) shared market data: reasoning is now the default across frontier models; DeepSeek R2 holds market share despite no new major release; open-source models represent 35–40% of tokens but a disproportionately small share of revenue; agentic workloads now exceed 15% of platform spend. Geoffrey Huntley (latent patterns) argued that software development now costs less than minimum wage ($10.
42/hour for Claude Code at API pricing), citing a K-shaped divide between lean model-first companies and slow incumbents. He cited a New Zealand founder who stopped backfilling 2.5 years ago and now has 20 people producing 30x the prior output.
Philip Kiely (Baseten) gave a technical deep-dive on quantization: legacy integer formats earned quantization its bad reputation, but modern floating-point formats (FP8, FP4, NVFP4 on Blackwell GPUs) are categorically better. NVFP4 uses block-wise scaling at n=16 with a secondary FP32 global factor. Real-world speed gains are 30–50% per precision step; weights are safe to quantize to 4-bit, but attention layers should not be touched.
Alisa Fortin and Guillaume Vernade (Google DeepMind) demoed VEO 3.1 (video), Nano Banana 2/Pro (image with Google Search grounding), and Lyria 3/Lyria Real Time (music). A 4K grid-image generation trick cuts batch image costs by up to 95%.
Anna Juchnicki (Intuit) detailed a governed LangGraph + Snowflake MCP agent that generates SQL and opens GitHub PRs without write access to production, using stateful re-entry for approvals and a self-review/repair loop. Kent C. Dodds demonstrated Cody, a personal MCP agent on Cloudflare Workers using Code Mode and Cloudflare artifacts, with secret-safe authenticated fetch so the LLM never sees credentials.
Rita Kozlov (Cloudflare) showed Code Mode cuts token usage 70% vs. vanilla MCP, and server-side Code Mode (search + execute) reduces context from 2M+ tokens to under 2,000. She calculated the agent infrastructure gap: 80–160 million CPUs needed at 50% concurrency for 8 billion personal agents, versus tens of millions produced per year — with isolates (100x more efficient than VMs) as the only viable path to scale.
Ben Davis closed by framing AI SDK evolution as three generations — Gen1 raw API wrappers, Gen2 ergonomic frameworks (Vercel AI SDK), Gen3 full coding agent SDKs (OpenCode, pi) — and argued Gen3 enables markdown-as-programs: natural language skill files coding agents execute with self-healing retry behavior.
Mhm. >> [music] [music] >> Ready Iman? >> [music] >> Woo! Good morning everyone. Hello Miami. How's everyone doing? >> Welcome to Miami. >> Yes. >> How's everybody doing? >> And in case you forgot where we are, uh there is a Q in my jersey. >> [laughter] >> Uh so we're bringing AI engineer to Miami today and I'm so grateful to see uh I can't really see that well because the light is so bright, [laughter] but I can kind of see your faces and I'm just so glad that you're all here to celebrate AI e...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...