101

Categorical Deep Learning: Moving AI from Alchemy to Science

Researchers propose Category Theory as a unifying mathematical framework for deep learning to move beyond empirical trial-and-error and enable neural networks to internalize algorithmic reasoning and structural logic.

102

César Hidalgo on The Infinite Alphabet and the Laws of Knowledge

César Hidalgo argues that knowledge is a non-fungible, collective phenomenon that follows physical-like laws of growth, diffusion, and decay, meaning it cannot be simply downloaded or copied without embodied experience.

103

AutoGrad and the Bayesian Brain: Dr. Jeff Beck on the Future of AI

Dr. Jeff Beck argues that true intelligence requires moving beyond function approximation and LLMs toward object-centered, Bayesian models grounded in macroscopic physics rather than language.

104

Why Every Brain Metaphor in History Has Been Wrong

The video explores how scientific simplifications and metaphors—from hydraulic pumps to computers—often harden into perceived realities, arguing that understanding requires recognizing the limits of these models.

105

Why AI Has a Plato Problem: Mazviita Chirimuuta on the Philosophy of Neuroscience

Professor Mazviita Chirimuuta argues that AI research often relies on a 'Platonic' assumption that the universe is written in mathematical code, ignoring the essential role of biological embodiment and active interaction in human cognition.

106

The Brain Is Just Specialized Agents Talking To Each Other — Dr. Jeff Beck

Dr. Jeff Beck discusses the mathematical foundations of agency, the mechanics of Energy-Based Models (EBMs), and a modular theory of intelligence where the brain is viewed as a collection of specialized agents.

107

Blaise Agüera y Arcas on Symbiogenesis and the Computational Nature of Life

Blaise Agüera y Arcas argues that life is embodied computation and that symbiogenesis—the fusion of replicators—is the primary engine of evolutionary novelty and complexity, rather than random mutation.

108

The Dangerous Illusion of AI Coding: Jeremy Howard on Software Engineering vs. Coding

Jeremy Howard argues that while AI excels at 'coding' as a style-transfer problem, it lacks the fundamental capacity for 'software engineering,' risking a future of 'understanding debt' and organizational enfeeblement.

109

Shinka Evolve: Open-Ended Program Search for Scientific Discovery

Robert Lange introduces Shinka Evolve, a sample-efficient framework that combines LLMs with evolutionary algorithms to automate program search and scientific discovery through the co-evolution of problems and solutions.

110

Measuring AI Progress: The METR Time Horizons Framework

Beth Barnes and David Rein of METR discuss their 'Time Horizons' methodology, which uses human completion time as a unified axis to measure AI capability and predict future progress.

111

Intelligence is Collective, Not Artificial: Prof. Michael I. Jordan's Perspective on AI and Economics

Professor Michael I. Jordan argues that intelligence is a collective economic system rather than a disembodied superintelligence, advocating for a shift from AI hype toward a multidisciplinary approach combining computer science, statistics, and economics.

112

Brad Carson on AI Weapons, Accountability, and the Fallacy of the AI Arms Race

Former Pentagon official Brad Carson argues that AI development is not an inevitable freight train and that the West can actively shape AI governance, prevent lethal autonomous weapons, and avoid a catastrophic arms race through strategic restraint and chip-level control.

113

Stanford CS336 Lecture 15: Mid-Training and Post-Training (SFT and RLHF)

This lecture explains the transition from base language models to instruction-following assistants through Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), emphasizing that data quality and curation are more critical than algorithmic complexity.

114

Stanford CS336 Lecture 16: Reinforcement Learning from Verifiable Rewards (RLVR)

This lecture explores Reinforcement Learning from Verifiable Rewards (RLVR), detailing how algorithms like GRPO replace complex value functions with group-based rewards to enable 'thinking models' capable of complex reasoning in math and coding.

115

Stanford CME296 Lecture 7: Evaluation of Text-to-Image Generation Models

This lecture outlines the methodologies for evaluating text-to-image models, distinguishing between aesthetics and prompt adherence, and detailing the transition from traditional mathematical metrics to MLLM-as-a-Judge frameworks.

116

Stanford CME296 Diffusion & Large Vision Models Lecture 8 Summary

A comprehensive review of image and video generation paradigms, covering the transition from diffusion and score matching to flow matching, and exploring the application of these techniques to video, image editing, and LLMs.

117

Stanford CS336 Lecture 17: Multimodal Language Modeling

This lecture explores the architecture of multimodal models, focusing on how vision-language models (VLMs) integrate image encoders like CLIP and SigLIP with large language models (LLMs) to achieve visual reasoning.

118

Stanford CS25: Serving Transformers - Lessons from the Trenches

Charles Frye of Modal discusses the engineering challenges of serving transformer models at scale, focusing on the critical distinction between prefill and decode phases and the optimization of hardware utilization.

119

Stanford CS25: Transformers United V6 - From Language Models to Native Multimodal Intelligence

Victoria Lin discusses the evolution of native multimodal language models, detailing how tokenization across modalities and specialized architectures like Mixture of Transformers (MoT) enable seamless integration of text, image, and audio.

120

Leveraging Geometry in Robot Learning: Stanford Robotics Seminar

Professor Robert Platt discusses how incorporating geometric structural priors and equivariance into robot learning models can significantly improve data efficiency and generalization over pose compared to generalist VLA models.

121

Stanford CS336 Language Modeling from Scratch: Inference Engines and Full-Stack Innovation

Guest lecturer Dan Fu discusses the critical role of inference engines and GPU kernels in transforming LLMs from mathematical objects into usable intelligence, introducing optimizations like Megakernels and the Parcae recurrent architecture.

122

Economics of the AI Supercycle: Baseten and the Shift to Custom Inference

Tuhin Srivastava, CEO of Baseten, argues that the AI economy is shifting from frontier models to custom, post-trained open-source models to achieve profitability, defensibility, and lower latency.

123

Sam Altman on Scale, AGI, and the Future of Frontier Systems

OpenAI CEO Sam Altman discusses the empirical power of scale, the evolution of ChatGPT and Codex, and the systemic risks of compute shortages and power concentration in the AI era.

124

AI in Healthcare: Cybersecurity Risks, Patient Empowerment, and Regulatory Sandboxes

Former U.S. Chief Data Scientist DJ Patil discusses the critical vulnerability of healthcare systems to AI-driven cyberattacks and the potential for AI to democratize health access through regulatory sandboxes and patient empowerment.

125

Stanford MS&E435 Economics of the AI Supercycle: Building AI Factories

Chase Lochmiller, CEO of Crusoe, discusses the massive capital expenditure required to build gigawatt-scale AI data centers, framing AI as a form of digital labor that drives GDP growth.

126

Must Haves For Agents in Production

To move LLM agents from demo to production, teams must implement seven critical controls: model control, prompt registries, guardrails, budget limiting, tool/MCP security, monitoring/tracing, and systematic evaluations.

127

The Era of Agents: Logan Kilpatrick on AI Studio and the Future of Building

Logan Kilpatrick discusses the evolution of AI Studio into a 'vibe coding' platform, the rise of agentic engineering, and Google's vision for a world where anyone can build software regardless of coding experience.

128

NVIDIA Nemotron 3 Nano Omni Release

NVIDIA has released Nemotron 3 Nano Omni, a compact, all-in-one multimodal model designed for agents that natively supports text, images, video, and audio in a single efficient architecture.

129

Claude Design Agentic Architecture: 6 Patterns for Vertical AI Agents

Claude Design utilizes a sophisticated stack of six agentic patterns—including context grounding, structured memory, and self-QA loops—to create a high-quality vertical agent that can be replicated for any industry.

130

IBM Granite Speech 4.1 Release: High-Throughput ASR Models

IBM has released Granite Speech 4.1, a suite of three 2B-parameter ASR models designed for edge deployment, offering specialized variants for accuracy, speaker diarization, and extreme throughput.

131

MiniCPM-V 4.6 release notes / what's new

MiniCPM-V 4.6 is a 1.3B parameter vision model designed for edge deployability and agentic workflows, featuring high token efficiency and flexible visual token compression.

132

OpenShell: Out-of-Process Enforcement for Secure LLM Agents

OpenShell provides a secure runtime for LLM agents by using a supervisor to enforce network, file, and credential policies outside the agent process, preventing prompt injection and jailbreak exploits.

133

Running Local AI on AMD Hardware

AMD's ROCm platform and Radeon AI Pro GPUs enable high-performance local execution of LLMs, image, and video generation models, offering a viable alternative to CUDA-based systems for privacy and cost-efficiency.

134

NVIDIA Nemotron 3 Ultra Release

NVIDIA has released Nemotron 3 Ultra, a 550B parameter Mixture-of-Experts model optimized for agentic workflows, featuring multi-teacher distillation and open training recipes.

135

NVIDIA Nemotron 3.5 ASR Release Notes

NVIDIA has released Nemotron 3.5 ASR, a 600-million parameter streaming speech-to-text model supporting 40 languages with cache-aware streaming for significantly reduced latency.

136

Anthropic Claude Fable 5 and Mythos 5 Release

Anthropic has launched Claude Fable 5, a safety-tuned Mythos class model that outperforms GPT-5.5 and Opus 4.8 in coding and legal benchmarks but introduces strict safety triggers and a mandatory 30-day data retention policy.

137

GLM 5.2 Release Notes and Performance Analysis

Z.AI has released GLM 5.2, an open-weights model that competes with frontier proprietary models in agentic coding and design, while offering significantly lower costs.

138

VibeThinker 3B: Scaling Reasoning in Small Language Models

VibeThinker 3B is a research model from Weibo AI Lab that uses reinforcement learning from verifiable rewards to match or beat models 300x its size on specific reasoning tasks like math and coding.