Categorical Deep Learning: Moving AI from Alchemy to Science
Researchers propose Category Theory as a unifying mathematical framework for deep learning to move beyond empirical trial-and-error and enable neural networks to internalize algorithmic reasoning and structural logic.
César Hidalgo on The Infinite Alphabet and the Laws of Knowledge
César Hidalgo argues that knowledge is a non-fungible, collective phenomenon that follows physical-like laws of growth, diffusion, and decay, meaning it cannot be simply downloaded or copied without embodied experience.
AutoGrad and the Bayesian Brain: Dr. Jeff Beck on the Future of AI
Dr. Jeff Beck argues that true intelligence requires moving beyond function approximation and LLMs toward object-centered, Bayesian models grounded in macroscopic physics rather than language.
Why Every Brain Metaphor in History Has Been Wrong
The video explores how scientific simplifications and metaphors—from hydraulic pumps to computers—often harden into perceived realities, arguing that understanding requires recognizing the limits of these models.
Why AI Has a Plato Problem: Mazviita Chirimuuta on the Philosophy of Neuroscience
Professor Mazviita Chirimuuta argues that AI research often relies on a 'Platonic' assumption that the universe is written in mathematical code, ignoring the essential role of biological embodiment and active interaction in human cognition.
The Brain Is Just Specialized Agents Talking To Each Other — Dr. Jeff Beck
Dr. Jeff Beck discusses the mathematical foundations of agency, the mechanics of Energy-Based Models (EBMs), and a modular theory of intelligence where the brain is viewed as a collection of specialized agents.
Blaise Agüera y Arcas on Symbiogenesis and the Computational Nature of Life
Blaise Agüera y Arcas argues that life is embodied computation and that symbiogenesis—the fusion of replicators—is the primary engine of evolutionary novelty and complexity, rather than random mutation.
The Dangerous Illusion of AI Coding: Jeremy Howard on Software Engineering vs. Coding
Jeremy Howard argues that while AI excels at 'coding' as a style-transfer problem, it lacks the fundamental capacity for 'software engineering,' risking a future of 'understanding debt' and organizational enfeeblement.
Shinka Evolve: Open-Ended Program Search for Scientific Discovery
Robert Lange introduces Shinka Evolve, a sample-efficient framework that combines LLMs with evolutionary algorithms to automate program search and scientific discovery through the co-evolution of problems and solutions.
Measuring AI Progress: The METR Time Horizons Framework
Beth Barnes and David Rein of METR discuss their 'Time Horizons' methodology, which uses human completion time as a unified axis to measure AI capability and predict future progress.
Intelligence is Collective, Not Artificial: Prof. Michael I. Jordan's Perspective on AI and Economics
Professor Michael I. Jordan argues that intelligence is a collective economic system rather than a disembodied superintelligence, advocating for a shift from AI hype toward a multidisciplinary approach combining computer science, statistics, and economics.
Brad Carson on AI Weapons, Accountability, and the Fallacy of the AI Arms Race
Former Pentagon official Brad Carson argues that AI development is not an inevitable freight train and that the West can actively shape AI governance, prevent lethal autonomous weapons, and avoid a catastrophic arms race through strategic restraint and chip-level control.
Stanford CS336 Lecture 15: Mid-Training and Post-Training (SFT and RLHF)
This lecture explains the transition from base language models to instruction-following assistants through Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), emphasizing that data quality and curation are more critical than algorithmic complexity.
Stanford CS336 Lecture 16: Reinforcement Learning from Verifiable Rewards (RLVR)
This lecture explores Reinforcement Learning from Verifiable Rewards (RLVR), detailing how algorithms like GRPO replace complex value functions with group-based rewards to enable 'thinking models' capable of complex reasoning in math and coding.
Stanford CME296 Lecture 7: Evaluation of Text-to-Image Generation Models
This lecture outlines the methodologies for evaluating text-to-image models, distinguishing between aesthetics and prompt adherence, and detailing the transition from traditional mathematical metrics to MLLM-as-a-Judge frameworks.
Stanford CME296 Diffusion & Large Vision Models Lecture 8 Summary
A comprehensive review of image and video generation paradigms, covering the transition from diffusion and score matching to flow matching, and exploring the application of these techniques to video, image editing, and LLMs.
Stanford CS336 Lecture 17: Multimodal Language Modeling
This lecture explores the architecture of multimodal models, focusing on how vision-language models (VLMs) integrate image encoders like CLIP and SigLIP with large language models (LLMs) to achieve visual reasoning.
Stanford CS25: Serving Transformers - Lessons from the Trenches
Charles Frye of Modal discusses the engineering challenges of serving transformer models at scale, focusing on the critical distinction between prefill and decode phases and the optimization of hardware utilization.
Stanford CS25: Transformers United V6 - From Language Models to Native Multimodal Intelligence
Victoria Lin discusses the evolution of native multimodal language models, detailing how tokenization across modalities and specialized architectures like Mixture of Transformers (MoT) enable seamless integration of text, image, and audio.
Leveraging Geometry in Robot Learning: Stanford Robotics Seminar
Professor Robert Platt discusses how incorporating geometric structural priors and equivariance into robot learning models can significantly improve data efficiency and generalization over pose compared to generalist VLA models.
Stanford CS336 Language Modeling from Scratch: Inference Engines and Full-Stack Innovation
Guest lecturer Dan Fu discusses the critical role of inference engines and GPU kernels in transforming LLMs from mathematical objects into usable intelligence, introducing optimizations like Megakernels and the Parcae recurrent architecture.
Economics of the AI Supercycle: Baseten and the Shift to Custom Inference
Tuhin Srivastava, CEO of Baseten, argues that the AI economy is shifting from frontier models to custom, post-trained open-source models to achieve profitability, defensibility, and lower latency.
Sam Altman on Scale, AGI, and the Future of Frontier Systems
OpenAI CEO Sam Altman discusses the empirical power of scale, the evolution of ChatGPT and Codex, and the systemic risks of compute shortages and power concentration in the AI era.
AI in Healthcare: Cybersecurity Risks, Patient Empowerment, and Regulatory Sandboxes
Former U.S. Chief Data Scientist DJ Patil discusses the critical vulnerability of healthcare systems to AI-driven cyberattacks and the potential for AI to democratize health access through regulatory sandboxes and patient empowerment.
Stanford MS&E435 Economics of the AI Supercycle: Building AI Factories
Chase Lochmiller, CEO of Crusoe, discusses the massive capital expenditure required to build gigawatt-scale AI data centers, framing AI as a form of digital labor that drives GDP growth.
Must Haves For Agents in Production
To move LLM agents from demo to production, teams must implement seven critical controls: model control, prompt registries, guardrails, budget limiting, tool/MCP security, monitoring/tracing, and systematic evaluations.
The Era of Agents: Logan Kilpatrick on AI Studio and the Future of Building
Logan Kilpatrick discusses the evolution of AI Studio into a 'vibe coding' platform, the rise of agentic engineering, and Google's vision for a world where anyone can build software regardless of coding experience.
NVIDIA Nemotron 3 Nano Omni Release
NVIDIA has released Nemotron 3 Nano Omni, a compact, all-in-one multimodal model designed for agents that natively supports text, images, video, and audio in a single efficient architecture.
Claude Design Agentic Architecture: 6 Patterns for Vertical AI Agents
Claude Design utilizes a sophisticated stack of six agentic patterns—including context grounding, structured memory, and self-QA loops—to create a high-quality vertical agent that can be replicated for any industry.
IBM Granite Speech 4.1 Release: High-Throughput ASR Models
IBM has released Granite Speech 4.1, a suite of three 2B-parameter ASR models designed for edge deployment, offering specialized variants for accuracy, speaker diarization, and extreme throughput.
MiniCPM-V 4.6 release notes / what's new
MiniCPM-V 4.6 is a 1.3B parameter vision model designed for edge deployability and agentic workflows, featuring high token efficiency and flexible visual token compression.
OpenShell: Out-of-Process Enforcement for Secure LLM Agents
OpenShell provides a secure runtime for LLM agents by using a supervisor to enforce network, file, and credential policies outside the agent process, preventing prompt injection and jailbreak exploits.
Running Local AI on AMD Hardware
AMD's ROCm platform and Radeon AI Pro GPUs enable high-performance local execution of LLMs, image, and video generation models, offering a viable alternative to CUDA-based systems for privacy and cost-efficiency.
NVIDIA Nemotron 3 Ultra Release
NVIDIA has released Nemotron 3 Ultra, a 550B parameter Mixture-of-Experts model optimized for agentic workflows, featuring multi-teacher distillation and open training recipes.
NVIDIA Nemotron 3.5 ASR Release Notes
NVIDIA has released Nemotron 3.5 ASR, a 600-million parameter streaming speech-to-text model supporting 40 languages with cache-aware streaming for significantly reduced latency.
Anthropic Claude Fable 5 and Mythos 5 Release
Anthropic has launched Claude Fable 5, a safety-tuned Mythos class model that outperforms GPT-5.5 and Opus 4.8 in coding and legal benchmarks but introduces strict safety triggers and a mandatory 30-day data retention policy.
GLM 5.2 Release Notes and Performance Analysis
Z.AI has released GLM 5.2, an open-weights model that competes with frontier proprietary models in agentic coding and design, while offering significantly lower costs.
VibeThinker 3B: Scaling Reasoning in Small Language Models
VibeThinker 3B is a research model from Weibo AI Lab that uses reinforcement learning from verifiable rewards to match or beat models 300x its size on specific reasoning tasks like math and coding.