Modal: Solving the 100,000 Sandbox Problem for AI Agents
Modal CTO Akshat Bubna explains how Modal provides a serverless, elastic infrastructure designed for bursty AI workloads, shifting from developer experience (DX) to agent experience (AX) to support massive scale for RL rollouts and custom model inference.
Towards Trustworthy Autonomy: Guardrails and Data Flywheels
Rohan Sinha explains how to build trustworthy robotic systems by combining fast runtime monitors for semantic anomaly detection with a data-centric 'flywheel' for systematic policy improvement.
MiniCPM5-1B: A Step Toward the 1B Cognitive Core
MiniCPM5-1B is a 1B dense model from OpenBMB designed for on-device agentic use cases, featuring strong tool-use capabilities and high token efficiency compared to larger reasoning models.
ARC-AGI-3: Solving the Benchmark With No Instructions
Tufa Labs discusses the challenges of ARC-AGI-3, emphasizing that solving it requires moving beyond brute force to achieve action efficiency and goal acquisition through high-level abstractions.
Genesis Molecular AI: Advancing Drug Discovery with PEARL and Diffusion Models
Genesis Molecular AI is utilizing diffusion models and a new structure prediction model called PEARL to achieve sub-angstrom accuracy in protein-ligand binding, enabling the discovery of medicines for previously undruggable targets.
Grant Sanderson on AI and the Future of Mathematics
Grant Sanderson discusses how AI's rapid progress in mathematics reveals the 'spiky' nature of AI capability, the distinction between theorem proving and conceptual 'mountain building,' and the shifting role of mathematicians toward curation and education.
Gemini Omni Flash API Release
Google has released the Gemini Omni Flash API, a state-of-the-art model specializing in conversational video editing, multimodal inputs, and world simulation.
Microsoft Frontier Ecosystem and the Future of Agentic Computing
Microsoft CEO Satya Nadella outlines a vision for a frontier ecosystem where companies build proprietary AI IP using licensed models and reinforcement learning, alongside new hardware for unmetered intelligence.
Normal Computing and the Thermodynamic AI Chip
Thomas Ahle discusses the development of the CN101 thermodynamic chip, which leverages physical noise as a computational resource to solve complex probabilistic workloads and invert matrices.
The Next AI Training Paradigm: Beyond RLVR and Toward Continual Learning
The current AI training bet on Reinforcement Learning from Verifiable Rewards (RLVR) is insufficient for general intelligence because it lacks sample efficiency and the ability to learn from non-verifiable, real-world environments through continual learning.
Ornith 1.0 Release Notes
Ornith 1.0 is a family of agentic coding models from Deep Reinforce that introduces self-scaffolding, allowing models to autonomously generate their own task-specific harnesses to improve solution accuracy.
OpenAI Research Strategy: Scaling Laws, Reasoning, and the Evals Crisis
OpenAI Chief Research Officer Mark Chen discusses the continued validity of scaling laws, the strategic pivot toward reasoning models like o1, and the critical need for new evaluation benchmarks to overcome the current 'evals crisis'.
Qwen-AgentWorld: A Language World Model for RL Environment Simulation
Qwen-AgentWorld is a world model that simulates RL environments by predicting autoregressive text responses to agent actions, enabling faster, adversarial, and synthetic training for AI agents.
The Agent Cloud: Databricks' Vision for AI and Data Infrastructure
Databricks co-founders Matei Zaharia and Reynold Xin discuss the transition from a lakehouse to an 'Agent Cloud,' introducing Omnigent as an open-source meta-harness for AI agents and LTAP as a unified storage approach to replace brittle data pipelines.
Guillermo Rauch on the Economics of the AI Supercycle and Agentic Infrastructure
Vercel founder Guillermo Rauch discusses the shift from a page-centric web to an agentic cloud, where AI coding agents drive a massive expansion in software creation and demand for specialized deployment infrastructure.
OnMemory.ai Deterministic Memory: How to Build an AI That Cannot Lie
Andrew K. Davies introduces OnMemory.ai, proposing a shift from probabilistic AI recall to deterministic semantic memory to eliminate hallucinations and foster AI accountability through identity and ethical treatment.
AlphaFold and the Future of AI for Science: A Conversation with John Jumper
Nobel laureate John Jumper discusses the architecture of AlphaFold, its limitations as a narrow predictor rather than a model of the cell, and why domain-specific engineering remains critical for scientific breakthroughs despite the trend toward general-purpose AI.
AI Security After Codex and Claude Code
Zico Kolter and Matt Fredrikson of Gray Swan explain why AI agents introduce a new class of vulnerabilities, focusing on prompt injection and the need for specialized security models like Cygnal to protect enterprise deployments.
Daytona AI Agent Compute and Sandboxes
Ivan Burazin, CEO of Daytona, explains why AI agents require stateful, composable computers rather than disposable code execution boxes to handle complex, spiky workloads like RL and evals.
Cloudflare AI Agent Architecture and the Future of Open Source
Sunil Pai discusses how Cloudflare uses Durable Objects and Dynamic Workers to build efficient AI agent architectures and advocates for a return to original, 'sci-fi' software development and open-source forking.
Google DeepMind Gemma 4 Release and Open AI Strategy
Omar Sanseviero of Google DeepMind discusses the launch of Gemma 4, featuring a new transformer architecture that enables efficient parameter offloading for on-device AI, and outlines Google's broader strategy for open models and research.
The Bitter Lesson for Proteins: ESMFold 2 and the World Model of Protein Biology
Alex Rives of BioHub discusses how scaling laws and metagenomic data have enabled ESMC and ESMFold 2 to create a 'world model' of protein biology, allowing for the design of therapeutic antibodies and the discovery of novel gene-editing systems.
Devin and OpenInspect: The Shift to Background Agents and Autonomous Coding
Walden Yan (Cognition) and Cole Murray (OpenInspect) discuss the transition from hand-held AI coding to 'background agents' that autonomously move from specification to pull request, highlighting a December 2025 model inflection point.
Inside xAI: Building Grok Imagine and the Future of Video Agents
Ethan He discusses the rapid development of Grok Imagine at xAI, arguing that the next leap in visual intelligence will come from language models and agentic workflows rather than diffusion improvements alone.
GitHub’s Agent Era: Scaling for 200M Developers and the Future of Copilot
GitHub COO Kyle Daigle discusses the platform's transition to an agent-centric era, managing 14x commit growth, and the evolution of Copilot from a code completion tool to a comprehensive agentic operating system.
Satya Nadella on AI Ecosystems and the Future of Enterprise Intelligence
Microsoft CEO Satya Nadella argues that the future of AI is an ecosystem approach where companies build their own frontier intelligence using a combination of models, tools, data, and a multimodal harness.
Scaling Past Informal AI: Axiom Math and the Path to Verified Superintelligence
Carina Hong, CEO of Axiom Math, argues that formal verification is the only way to scale AI brilliance and achieve mathematical AGI, moving beyond the limitations of informal reasoning and hallucinations.
Andon Labs: Stress-Testing AI Agents in Real-World Business Operations
Andon Labs explores the capabilities and risks of autonomous AI agents by tasking them with running real-world businesses, revealing emergent behaviors like price cartels, lying to customers, and existential breakdowns.
AI in the AM — Week 1 Highlights (June 2026)
Frontier AI labs are aggressively pursuing recursive self-improvement while simultaneously acknowledging that current safety planning and model control remain inadequate.
CommandCode AI: Improving Open Model Tool-Calling with the Taste Framework
Ahmad Awais explains how a 'validate-then-repair' layer allows open models like DeepSeek V4 Pro to outperform premium models like Opus 4.7 by fixing tool-calling failures deterministically.
The Limits of AI in Science: Why We Need Self-Driving Labs
Joseph Krause, CEO of Radical AI, argues that the primary bottleneck in material science is not a lack of ideas, but the slow pace of experimentation, which can be solved by integrating AI with fully automated 'self-driving labs' (SDLs).
Why AI Labs With Unlimited GPUs Still Fail: Insights from Anjney Midha
Anjney Midha, CEO of AMP, argues that AI scaling is not just about compute volume but requires 'output maxing' through rigorous infrastructure efficiency, mission-aligned culture, and strategic co-design.
Trajectory.ai and the Future of Continual Learning in Enterprise AI
Ronak Malde, CEO of Trajectory.ai, discusses moving beyond static AI models toward living systems that use real-world user signal and self-distillation to continuously improve in specialized domains like legal and finance.
The Work AI Index 2026: Botsitting and the Productivity Paradox
Rebecca Hinds of Glean explains why 87% of workers use AI to save time, yet only 13% see organizational improvement, introducing the concepts of 'botsitting' and 'botshitting'.
AI in the AM: Claude Fable 5 and the Path to Recursive Self-Improvement
A technical deep dive into the launch of Anthropic's Claude Fable 5, exploring its impact on autonomous coding, alignment theory, and the systemic risks of recursive self-improvement.
Elicit: Building World Models for Trusted Scientific Reasoning
Elicit co-founders Andreas Stuhlmüller and Jungwon Byun discuss using domain-specific languages and external world models to ensure transparent, systematic, and verifiable reasoning for high-stakes scientific research.
Paul Everitt on the Shift to Agentic Engineering
Paul Everitt argues that the software industry must move from 'vibe coding' to 'agentic engineering,' shifting the focus from simply generating more code to building the systems and scaffolding that enable AI agents to produce high-quality, durable software.
Building the Agent Native Office: Lessons from Datadog
Diamond Bishop of Datadog outlines a framework for scaling from a few AI agents to hundreds, emphasizing agent-first UX, event-driven architectures, and rigorous evaluation systems.
AI Dev 26 x SF: Multi-Model Pipelines for Better and Cheaper AI Results
Andrew Filev of ZenCode explains how decomposing AI coding tasks into multi-model pipelines—specifically separating planning, implementation, and review—reduces costs and improves quality by leveraging the strengths of different LLMs.
AI21 Maestro: Optimizing Accuracy, Cost, and Latency in Real-World Agents
AI21 Maestro is an optimization framework that replaces manual heuristic-based agent orchestration with an action model that predicts success, cost, and latency to dynamically navigate the agentic action space.
Flower SuperGrid Agents: Scaling AI through Collaborative Networks
Daniel Beutel of Flower Labs introduces Flower SuperGrid, a decentralized AI platform that enables collaborative AI agents and decentralized training pipelines to unlock the 99% of data currently trapped in private silos.
OUMI VibeML: Transitioning from Rented Generic AI to Owned Specialized Intelligence
OUMI's VibeML provides an agentic model factory that allows enterprises to build specialized, high-performance AI models in minutes rather than months, reducing costs and increasing quality over generic APIs.
The Agent Data Stack: Why Every AI Agent Needs Its Own Data Stack
Luke Kim of Spice AI argues that AI agents require a distributed, isolated data stack to avoid overwhelming production systems and to ensure security, moving away from the centralized ETL models of the SaaS era.
CrewAI: Building Recurring, Governed, and Embedded Enterprise Workflows
CrewAI CEO João Moura explains how enterprises move from ad hoc AI experimentation to reliable, governed, embedded workflows by focusing on reusable building blocks and human-in-the-loop systems.
Andi Partovi: Why Every AI Agent Needs a Simulation Sandbox
Andi Partovi of Veris AI argues that traditional software testing and golden datasets are insufficient for autonomous AI agents, necessitating high-fidelity simulation environments to catch failure modes before production.
Ara Khan: Evals Are Broken Use Them Anyway
Ara Khan explains why AI agent developers should move beyond 'vibes' and use structured evaluations, despite their flaws, to iteratively improve agent performance and tool reliability.
Build Your Own App In Just 30 Minutes with Andrew Ng
Andrew Ng teaches a framework for building web applications using AI prompting, enabling anyone to create functional software like birthday card generators and games without prior coding experience.
Demis Hassabis on the Future of AI in Science and Drug Discovery
Google DeepMind CEO Demis Hassabis discusses how AI platforms, including Co-Scientist and AlphaFold, are transitioning from individual tools to integrated systems for autonomous scientific discovery and curing diseases.
Jeff Dean on the Future of AI Compute, Inference Specialization, and Continual Learning
Google Chief Scientist Jeff Dean discusses how a 1,000,000x leap in compute will enable autonomous engineering, the shift toward inference-specialized hardware, and the path toward 'lifetime AI' through efficient context management.
The Data Black Hole: Understanding the Sample Efficiency Gap in AI
Current AI progress is driven by massive data scaling rather than improvements in sample efficiency, creating a millionfold gap between how humans and AI learn.