This hands-on workshop from AI Engineer Europe 2026 features three Google DeepMind engineers — Paige Bailey, Guillaume (Guom) Vernade, and Ian Valentine — demonstrating the breadth of Google's generative AI stack live in front of an audience, covering cloud APIs, on-device models, and a complete multimodal media pipeline.
Paige Bailey opens with a tour of AI Studio and recently released models. She demonstrates Gemini 3.1 Flash Light — priced at roughly $0.25 per million tokens, nearly an order of magnitude cheaper than Gemini 3.1 Pro — analyzing a dinosaur YouTube video frame-by-frame at about one frame per second with Google Search grounding enabled. She then uses AI Studio's Build feature (comparable to v0.dev or Lovable) to generate a full bookshelf-scanning app from a single prompt, complete with Google OAuth login, Firestore persistence, image upload, and Google Search grounding for book metadata — all without writing a line of code manually. The Build feature auto-generates Firestore security rules, surfaces a version history, supports GitHub sync, and allows custom API key injection. Paige also demos Gemini Live for real-time multimodal conversation with screen sharing and multilingual support, and Genie 3 — a world-model system composing Imagen (Nano Banana), VO, and Gemini — to render a playable 60-second pixel world featuring a pink sparkly squirrel on Regent's Canal, generated entirely as raw pixel frames with no game engine or 3D assets.
Guillaume Vernade leads a Jupyter notebook workshop using Wind in the Willows (Project Gutenberg) as source material to illustrate all four major gen media models end-to-end. He uses Gemini in chat mode with structured outputs to generate character-consistent image prompts, then feeds them to Imagen 3 (Nano Banana 2) — which supports multiple aspect ratios, image-grounded generation, and reverse image search — to produce character portraits and chapter illustrations in a dark-fantasy style. He passes the final illustrations as starting frames into VO 3.1 Light (5 cents per image, the cheapest video generation model in the stack) to animate each scene, and scores every chapter with LIA 3, Google's first publicly available music-generation-via-API model capable of generating 30-second clips or full 3-minute songs with lyrics. Guillaume notes that Gemini is trained on gen-media prompts internally, making it uniquely capable at prompt generation for these models. He also introduces LIA Realtime, a lesser-known model that generates music indefinitely in real time like a DJ responding to live prompts. A new service tier parameter — "flex" for asynchronous, cost-efficient requests, and "priority" for 2x cost but higher reliability — is highlighted as a practical cost optimization tool.
Ian Valentine closes with Gemma 4, released the prior Thursday under Apache 2.0. The Gemma 4 family includes: E2B and E4B effective models (designed for phones, Raspberry Pis, and Jetson Nanos, with per-layer embedded architecture that pages embeddings from flash); a 26B mixture-of-experts model with only 4B activated parameters (requiring ~22 GB RAM for full context); and a 31B dense flagship. Ian demos the 26B model running locally in LM Studio on an M4 Mac, served on an OpenAI-compatible endpoint on port 1234. He then launches 10 parallel sub-agents through a terminal orchestrator — all hitting the local Gemma 4 26B — generating SVGs simultaneously, illustrating multi-agent workflows with no cloud API involved. Finally, using open-code (configured with a single JSON file pointing at localhost), Ian has the local 31B model read a game spec, build a working space-shooter called Nebula Drift including debugging its own syntax errors across multiple tool-call iterations, and one-shot recreate a web page from a screenshot. All Gemma 4 models are multimodal (image and video; audio on E2B/E4B) and support function calling, thinking mode, and agent skills natively.
My name is Paige. I started doing machine learning a long time ago. Um around 2009 2010. Um was primarily working with um though it feels like forever ago. I was just talking with a friend about this recently. Um, back in 2009 2010, it was kind of wild that companies would even trust open source software to do business critical work. Um, and so I was contributing to things like numpy, sci-fi, like little like little antenna these microphones. Um, >> yeah. >> Sure. Cool. Cool. Um so numpy, scypi,...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...