Sign InOpen Brain
Atlas / Model

Generative Media

Open JSONConfidence: Auto-collectedLast updated 2026-07-30

Current Answer

No editorial synthesis yet — the evidence below is collected automatically from source labels. A current answer lands here once an editor approves one.

Evidence

MiniMax H3 now available on AI Gateway
Vercel · 2026-07-30

MiniMax H3 brings short 2K video generation to Vercel AI Gateway, with text, keyframe, and multimodal reference inputs. Reference and keyframe modes cannot be combined.

Grok Voice Think Fast 2.0 now available on AI Gateway
Vercel · 2026-07-29

Grok Voice Think Fast 2.0 brings speech-to-speech reasoning and earlier tool calls to Vercel’s realtime API, with server-minted tokens keeping gateway keys off clients.

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
arXiv · 2026-07-27

ModernMOE applies efficient expert-routing patterns from LLMs to diffusion transformers, improving convergence and quality-cost balance without relying only on larger parameter counts.

Evaling Video Slop — Maor Bril, Character.ai
AI Engineer · 2026-07-25

Video evaluators can reward polish while missing frozen action, broken physics, or failed storytelling. Builders need time-aware criteria and human-calibrated data, not frame quality alone.

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber
AI Engineer · 2026-07-24

Uber’s image-editing agent uses routing, iterative QA, golden-set gates, and production feedback to avoid costly edits, hallucinated food, and quality regressions.

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
arXiv · 2026-07-23

MedGame turns static clinical cases into executable decision stories with separate narrative and orchestration stages, a useful architecture pattern for case-grounded learning agents.

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
arXiv · 2026-07-14

A frozen diffusion language model can transcribe speech by refining the full transcript in parallel. The prototype trains a small audio interface and reaches 6.6% WER in roughly eight steps.

Seedream 5.0 Pro is now available on AI Gateway

Seedream 5.0 Pro adds image generation and editing to Vercel AI Gateway, targeting reliable text rendering and dense infographic layouts through the AI SDK.

Evidence-Backed Video Question Answering

E-VQA requires video answers to include tracked pixel-level evidence, revealing when good QA scores hide weak perception and supplying grounded training data.

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

SearchGen-Bench shows open image generators score 21–28/100 on long-tail entities, and naive search retrieval only adds noise; a teach-then-search co-training recipe learns when to retrieve versus rely on weights.

Meta 3D AssetGen: Generating 3D Worlds With AI

Meta's Tech Podcast covers AssetGen, its foundation model for generating 3D assets from text, and the path toward AI-generated worlds in Horizon Studio. A podcast episode, so light on specifics.

Start building with Nano Banana 2 Lite and Gemini Omni Flash

Two new Gemini API models: Nano Banana 2 Lite generates 1K images in ~4s at $0.034 each, and Omni Flash does video at $0.10/sec in public preview — cheap enough to wire asset generation into agent pipelines.

Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models

IAAN boosts selected audio-encoder neurons at inference, improving fine-grained speech perception across three models without retraining or labels.

The latest AI news we announced in June 2026

Google's June roundup: Gemma 4 12B runs locally in 16GB of memory, Gemini 3.5 Flash adds computer use for desktop, mobile, and browser agents, and Nano Banana 2 Lite ships as a cheaper image model.

Stable permalink · evidence auto-collected from source labels · synthesis maintained by feed7 editorial