AI Engineer Miami Day 2 (April 21, 2026) featured eight back-to-back presentations covering agentic coding adoption, inference hardware breakthroughs, on-device generative AI, specialized sub-agent models, context engineering, agentic memory, MCP vs. CLI evals, open-source model flexibility, and IDE redesign.
David House (G2i) opened with case studies on agentic coding adoption. He argued frameworks should constrain beginners while amplifying experts, chaining a /brief skill (PM persona), /spec skill (architect persona), /code or /tdd, /review, and /draft-PR skill. His thesis: agentic use is not intuitive for any engineer — the 'slop cannon vs. skill issue' binary is false; proper frameworks bridge that gap.
Sarah Chieng (Cerebras) diagnosed 'latency debt': generation speed has stayed flat at 50–150 tokens/second across the Claude and Gemini families while context and output volumes tripled. An OpenRouter/a16z study across 100 trillion real-world tokens showed input token lengths grew 4x and output tokens 3x in one year alone. The Cerebras/OpenAI Codex Spark model runs at 1,200 tokens/second — 20x faster — using the Wafer Scale Engine 3 with on-chip distributed SRAM to eliminate off-chip memory-wall bottlenecks. Nvidia's $20B Groq acquisition and Cerebras–AWS disaggregated prefill/decode partnership are cited as macro validation.
Lech Kalinowski (CallStack) demonstrated fully local latent diffusion on an Android NPU via ONNX Runtime, replacing text prompts with ambient light sensor values mapped to a latent vector — zero cloud API calls, 8–24-step denoising on-device.
Tejas Bhakta (Morph LLM) framed 'Software 3.5' as agents prompting other agents. Morph builds specialized sub-agent models for non-frontier tasks: a code-search model supporting 12 parallel tool calls (80k input/~200 output tokens), a context compaction model at 33,000 tokens/second, and a fast-apply diff model. Cross-customer data shows doubling speed without hurting accuracy roughly doubles conversion rates; the code-search sub-agent yields a 3% SWE-bench Pro improvement.
Rick Blalock (Agentuity) argued coding agents are now a universal software primitive — building chatbots, RAG pipelines, and orchestrating sub-agents. He traced the arc from AutoGPT hype through brittle LangChain/CrewAI orchestration to present convergence, noting existing cloud platforms remain designed for stateless web workloads and are not purpose-built for long-running agents.
Nyah Macklin (Neo4j) presented context graphs for auditable AI. Peer-reviewed results showed accuracy improving from 37% (base) to 54% (fine-tuning) to 91% (graph RAG + knowledge graph) on a financial-services benchmark. Context graphs capture decision traces and causal chains behind every agent decision — essential for regulated domains like credit approval where agents must be explainable and auditable.
Laurie Voss (Arize AI) ran 500 structured evaluations comparing GitHub's official MCP server, a verbose 2,187-line GH skill (LoHub), and a short opinionated GH skill (Claude Skills Vault) using Claude Opus 4.6 via the Claude Agent SDK. All three scored correctness in the high 80s with 100% on read-only tasks. MCP used 12 tool calls vs. 5 per task on complex queries and incurred 6x higher cost — verbose JSON payloads overflowed context windows, forcing the agent to use bash/jq. Voss concluded the debate is a false dichotomy: CLI/skills win for tools with strong training-data coverage; MCP wins for remote, proprietary, OAuth-gated tools and consumer contexts. Production agents (Cursor, Claude Code) use both.
David Gomes (Cursor) presented Cursor 3.0, built from scratch outside VS Code's architecture. His own tab usage fell from 1,400 per month (September 2025) to near-zero as agents took over. The redesign prioritizes video playback for verifying agent output, canvas data-science views, and DAG-style orchestration UIs. Stefan Avram (OpenCode) noted enterprises cycle through Resist → Rush → Reign phases, and Ramp's 'Inspect' agent — built on OpenCode's open codebase — now authors 30% of Ramp's merged PRs.
Hello. Hello. Good morning. How's everybody doing? >> Whoa. Okay, that's the energy I'm asking for. >> No way. You came back. Well, I also know some people uh just came here as their first day. So, uh welcome back everybody. And for people who are new, welcome to AI Engineering Miami day two. >> Okay, some quick questions. Who learned something awesome yesterday that they can't wait to try very soon? >> Oo, I see hands. >> Okay, who made some new LinkedIn connections? >> Oh, I see hands. >> Who ...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...