Sign InOpen Brain
Atlas / Model

Model Selection

Open JSONConfidence: Auto-collectedLast updated 2026-09-10

Current Answer

No editorial synthesis yet — the evidence below is collected automatically from source labels. A current answer lands here once an editor approves one.

Evidence

AI EngineerWorkshopTranscript Verified
Building eval sets that survive model swaps — AI Engineer workshop
Eval sets usually die when you change models. This workshop shows how to write ones that transfer.
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
arXiv · 2026-09-10

MoE models overfit repeated training data earlier than dense models, with total parameter count driving the effect. Strong masking helps, but unique data remains the stronger baseline.

Can Edge-Deployable Vision-Language Models Identify Species?
arXiv · 2026-09-10

For edge vision, specialized training outweighed model size: 300M-parameter BioCLIP beat 2–8B VLMs, while field imagery degraded every model and open-set prompts produced invented species.

GPT-6 Astra: The next generation in intelligence for work
OpenAI · 2026-09-09

GPT-6 Astra combines reasoning, computer use, and writing and design judgment, potentially widening the range of knowledge-work tasks one agent can handle.

DeepSeek V4.1 Flash now available on AI Gateway
Vercel · 2026-09-09

DeepSeek V4.1 Flash brings vision, tool use, reasoning, and prompt caching to Vercel AI Gateway, with direct setup paths for Claude Code, Codex, and Cursor.

GPT Image 2.5 Flare and Sunburst now available on AI Gateway
Vercel · 2026-09-08

Vercel AI Gateway adds GPT Image 2.5 Flare for faster iteration and Sunburst for tighter composition control. Both support generation, reference-based edits, and transparent backgrounds.

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
arXiv · 2026-09-04

Molecular benchmark accuracy can reflect retrieval of published values rather than prediction. Contamination audits should test digit-level recall and repeat runs across reasoning settings.

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
arXiv · 2026-09-04

A compact classifier distilled from LLM rationales handled product-pair decisions without inference-time LLM calls. Category adapters improved accuracy further while retaining large speed and cost gains.

Ling 3.0 Flash Sante is now available on AI Gateway for free
Vercel · 2026-09-04

Vercel AI Gateway now serves the medical-focused Ling 3.0 Flash Sante free through October 4. Use its free-only model ID to prevent requests from converting to paid usage afterward.

Rethinking On-Policy Distillation of Large Language Models II: One Training Example
arXiv · 2026-09-03

On-policy distillation recovered most full-data gains from one query; diverse rollouts mattered more than dataset size, while slow student alignment remained the bottleneck.

Legora reviewed 41 documents in minutes with GPT-6 Astra
OpenAI · 2026-09-03

Legora says GPT-6 Astra reviewed 41 documents within minutes, caught every planted error, and improved its workflow result by nearly 40%, though the underlying measure is unspecified.

Playco cut manual fixes 50% prototyping games with GPT-6 Astra
OpenAI · 2026-09-03

Playco reports that GPT-6 Astra halved manual fixes while producing three themed game prototypes from one grey-box base, suggesting less cleanup in model-driven iteration.

GPT-6 Astra: A new generation of intelligence
OpenAI · 2026-09-03

OpenAI introduces GPT-6 Astra with claimed advances in computer use, coding, cybersecurity, and science. The supplied material gives no benchmarks or implementation details to assess those gains.

Safety overview: GPT-6 Astra
OpenAI · 2026-09-03

GPT-6 Astra is OpenAI's first model rated Critical for cybersecurity capability under its Preparedness Framework, a material consideration for security-sensitive agent access and controls.

Post-Training Language Models for Gold-Medal Performance in Coding Competitions
arXiv · 2026-09-02

A coding-specialized model paired post-training with an iterative generate-evaluate-refine loop to exceed the top IOI 2026 human score. The reusable idea is feedback-driven test-time search.

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
arXiv · 2026-09-02

A UE5M3 block-scaling recipe trained an 8B model in FP4 without Hadamard transforms or BF16 final layers, while reporting better losses and downstream estimates than the compared recipe.

Muse Spark 1.3 now available on AI Gateway
Vercel · 2026-09-02

Muse Spark 1.3 gives coding agents a 1M-token, multimodal model through Vercel, with a cheaper contributor tier that permits Meta to train on submitted inputs and outputs.

Gemini 3.8 Flash now available on AI Gateway
Vercel · 2026-09-02

Gemini 3.8 Flash brings multimodal input, tool calling, web search, and default reasoning to coding agents through Vercel. Its temporary 50% discount runs through December 31.

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
arXiv · 2026-09-01

When quantizing an open model, spend a small extra precision budget across the network before protecting a few “important” layers; causal tests found the damage was usually diffuse.

Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
arXiv · 2026-09-01

Small proxy models may be enough to choose an SFT-versus-RL annotation split: the paper finds broad near-optimal ranges that transfer to larger models.

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
arXiv · 2026-09-01

A production-derived post-training recipe consolidated more than 200 internal apps onto one self-hosted model by training separate experts for distinct quality gaps, then merging them.

Claude Fable 5.1 now available on AI Gateway
Vercel · 2026-09-01

Claude Fable 5.1 reaches Vercel AI Gateway with ordered fallbacks for classifier refusals, but its 30-day retention policy rules out zero-data-retention workloads.

Qwen 3.8 Max 0902 now available on AI Gateway
Vercel · 2026-09-01

Vercel’s gateway now exposes a pinned Qwen snapshot aimed at larger coding projects and longer agent runs, with a dated model ID that prevents silent upgrades.

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning
arXiv · 2026-08-31

A controlled ontology-learning study finds model size is a weak selector on its own. Dense 27B models beat larger sparse models on one task, while MoE models led another.

Agentic Sites: Building Hyper Personalized Websites — Carlos Sanchez, Adobe
AI Engineer · 2026-08-29

Adobe’s prototype assembles intent-specific page blocks from existing site content in roughly a second, making model latency and per-site evaluation part of frontend architecture.

The Half Life of Agent Infrastructure — Ben Kus, Box
AI Engineer · 2026-08-29

Agent architectures are expiring quickly. Keep model, search, and orchestration choices replaceable, and evaluate platforms by how well they handle repeated change.

Hy4 Preview now available on AI Gateway
Vercel · 2026-08-28

Tencent’s Hy4 Preview is now callable through Vercel AI Gateway and selectable in coding agents, adding an open MoE option with a 1M-token context window.

Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
arXiv · 2026-08-27

Three ways to combine RLVR domain experts perform similarly on average but diverge by task. Choose Merge for cheap reuse, Mix RL for training from pooled data, and MOPD for preserving expert gains.

Ling 3.0 Flash Fin now available on AI Gateway for free
Vercel · 2026-08-27

Ling 3.0 Flash Fin adds a finance-focused reasoning and tool-calling option to AI Gateway, with separate model IDs for automatic billing or a hard stop after the free period.

Qwen 3.8 Flash now available on AI Gateway
Vercel · 2026-08-26

Qwen 3.8 Flash is now selectable in Vercel AI Gateway and coding agents, with text-and-image input, a 1M-token context window, and responses up to 65k tokens.

GLM 5.3 Flash now available on AI Gateway
Vercel · 2026-08-26

GLM 5.3 Flash joins Vercel AI Gateway with text and vision input, a 1M-token context window, function calling, structured output, and streaming.

Wan 3.0 now available on AI Gateway
Vercel · 2026-08-25

Wan 3.0 gives AI Gateway one video model ID for text, image, frame, and reference workflows, with async renders up to 30 seconds at 1080p and synchronized audio.

MiniMax M3 and M2.7 are free on AI Gateway
Vercel · 2026-08-25

AI Gateway offers temporary free routes for MiniMax M3 and M2.7, but the -free model IDs hard-fail after September 6; standard IDs preserve provider fallback at normal rates.

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
AI Engineer · 2026-08-22

Per-task model routing cut the demonstrated coding session’s cost from 44¢ to 14¢ with similar completion time, but builders still need workload-specific evals to validate quality.

Asymmetric Capacity Allocation in Self-Refinement Pipelines
arXiv · 2026-08-21

Self-refinement pipelines need not use equally capable models: invest capacity in generation and revision, while a small critic may preserve gains at lower compute cost.

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
arXiv · 2026-08-21

CLEAR conditionally activates a safety adapter instead of applying safety tuning to every prompt, reducing harmful completions while limiting benign-task degradation.

Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End
AI Engineer · 2026-08-20

Lease End replaced a fine-tuned intent classifier with skills and runtime context, cutting production fixes from about a week to under an hour. Higher API spend was offset by lower maintenance cost.

TokEval: A Tokenizer Evaluation Suite
arXiv · 2026-08-18

TokEval links tokenizer properties to language, math, and code performance, offering cheaper screening signals before committing compute to pretraining sweeps.

Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry
arXiv · 2026-08-18

A single-GPU SBERT beat the reported zero-shot LLM and vendor baseline for invoice coding, suggesting narrow, private classifiers can outperform broader models with modest local data.

GLM 5.3 now available on AI Gateway
Vercel · 2026-08-18

GLM 5.3 is available through Vercel AI Gateway for coding agents, retaining a 1M-token context window while claiming better long-horizon engineering with fewer output tokens.

GPT-5.6 Sol is 50% off on AI Gateway for the next month
Vercel · 2026-08-17

Vercel cut GPT-5.6 Sol pricing by 50% through September 18, making direct AI Gateway runs cheaper across every tier without changing model IDs or agent configs.

Stable permalink · evidence auto-collected from source labels · synthesis maintained by feed7 editorial