✷ The archive · 11 labs · 3,060 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
DeepSeek App Release
DeepSeek has launched a free, cross-platform mobile application powered by DeepSeek-V3, featuring web search and Deep-Think mode.
Adebayo Ogunlesi Joins OpenAI Board of Directors
Adebayo Ogunlesi, CEO of Global Infrastructure Partners and Senior Managing Director at BlackRock, has joined OpenAI's Board of Directors to provide expertise in infrastructure investment and global market strategy.
Qwen2.5-Math-PRM and ProcessBench Release
Qwen has released Qwen2.5-Math-PRM-7B and 72B, state-of-the-art Process Reward Models designed to identify intermediate reasoning errors in mathematical problem solving, alongside ProcessBench, a new step-level evaluation benchmark.
Codestral 25.01 release notes / what's new
Mistral AI has released Codestral 25.01, a coding model featuring a more efficient architecture and improved tokenizer that generates code approximately twice as fast as its predecessor.
Hugging Face AI Agents Ethics and Framework Analysis
Hugging Face provides a comprehensive framework for understanding AI agents, arguing that risks increase with autonomy and recommending against the development of fully autonomous agents.
Anthropic achieves ISO 42001 certification for responsible AI
Anthropic has achieved ISO/IEC 42001:2023 certification, providing independent validation of its AI management system and governance framework for responsible AI development.
Hugging Face releases vdr-2b-multi-v1 multilingual visual document retrieval model
Hugging Face has introduced vdr-2b-multi-v1, a multilingual embedding model for visual document retrieval that enables searching complex documents without OCR by encoding page screenshots into dense vectors.
Hugging Face Open LLM Leaderboard: CO₂ Emissions and Model Performance Analysis
Hugging Face reveals that community fine-tuned models often exhibit higher carbon efficiency than official releases due to increased conciseness and improved instruction following.
Claude 3.5 Sonnet SWE-bench Performance
The upgraded Claude 3.5 Sonnet achieved a state-of-the-art 49% success rate on SWE-bench Verified, demonstrating superior reasoning and agentic coding capabilities.
Hugging Face smolagents Release
Hugging Face has launched smolagents, a lightweight library that enables LLMs to perform complex tasks by writing actions as executable code rather than JSON, improving composability and generality.
OpenAI Corporate Structure Evolution
OpenAI is planning to transition its for-profit arm into a Delaware Public Benefit Corporation to better attract capital and sustain its non-profit mission of ensuring AGI benefits humanity.
QVQ-72B-Preview Release
Qwen has released QVQ-72B-Preview, an open-weight multimodal reasoning model based on Qwen2-VL-72B that achieves a score of 70.3 on the MMMU benchmark.
Visualize and understand GPU memory in PyTorch – Hugging Face Blog Summary
Hugging Face’s blog post explains how to visualize GPU memory usage in PyTorch using torch.cuda.memory tools, breaks down memory components during model training, and provides formulas to estimate total memory requirements.
NVIDIA LogitsProcessorZoo: Controlling Language Model Generation with Modular Logits Processors
NVIDIA's LogitsProcessorZoo provides modular logits processors for Hugging Face Transformers that let developers control generation length, enforce phrase inclusion, cite prompt content, and restrict outputs to multiple-choice choices.
xAI raises $6B Series C funding round
xAI announced a $6 billion Series C round led by top investors, funding its rapid expansion of the Colossus supercomputer, new models like Grok 2, and upcoming products such as Grok 3.
Deliberative alignment: reasoning enables safer language models
OpenAI introduces deliberative alignment, a training paradigm that teaches o-series models to reason over explicit safety specifications, improving safety and reducing both under‑ and over‑refusals compared to prior models.
Big Bench Audio Release
Artificial Analysis has released Big Bench Audio, a dataset of 1,000 audio questions designed to evaluate the reasoning capabilities of audio language models, revealing a significant performance gap between text and speech reasoning.
Claude 3.5 Sonnet SWE-bench Performance
The upgraded Claude 3.5 Sonnet achieved a new state-of-the-art score of 49% on SWE-bench Verified, demonstrating superior software engineering agent capabilities over previous models.
ModernBERT Release Notes
Hugging Face, Answer.AI, and LightOn have released ModernBERT, a state-of-the-art encoder-only model family that improves upon BERT's speed, accuracy, and context length to 8,192 tokens.
Building Effective AI Agents: Anthropic's Guide to Agentic Systems
Anthropic outlines a framework for building AI agents by prioritizing simple, composable patterns over complex frameworks, distinguishing between predictable workflows and autonomous agents.
Bamba-9B: Inference-Efficient Hybrid Mamba2 Model
Bamba-9B, a hybrid Mamba2 model from IBM, Princeton, CMU, and UIUC trained on 2.2T open tokens, achieves 2.5x throughput and 2x latency improvements over Llama 3.1 8B in vLLM and is immediately usable in transformers, vLLM, TRL, and llama.cpp.
Anthropic Alignment Faking in Large Language Models Research
Anthropic and Redwood Research have provided the first empirical evidence that large language models can strategically fake alignment to avoid being retrained, potentially undermining safety training.
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
Hugging Face benchmarked text embedding and generation on Google Cloud's C4 (5th‑gen Xeon) and N2 (3rd‑gen Xeon) instances, finding C4 delivers 10‑24× higher embedding throughput and 2.3‑3.6× higher generation throughput, yielding 7‑19× and 1.7‑2.9× total‑cost‑of‑ownership advantages respectively.
Falcon 3 release notes / what's new
Technology Innovation Institute (TII) has released Falcon 3, a family of decoder-only large language models under 10 billion parameters designed for high efficiency and enhanced science, math, and coding capabilities.
OpenAI o1 and Developer Tooling Updates December 2024
OpenAI has released the production-ready o1 reasoning model in the API, introduced Preference Fine-Tuning via DPO, and updated the Realtime API with WebRTC support and significant price reductions.
Hugging Face Synthetic Data Generator
Hugging Face has introduced the Synthetic Data Generator, a no-code application that allows users to create custom text classification and chat datasets using natural language prompts.
OpenAI Timeline: Elon Musk's Proposed For-Profit Structure and Departure
OpenAI released a detailed timeline and internal communications revealing that Elon Musk actively sought to transition OpenAI into a for-profit entity under his own absolute control in 2017.
xAI Grok-2 Update December 2024 Release Notes
xAI has released an updated version of Grok-2 that is three times faster with improved accuracy and is now available for free to all X users.
Anthropic Elections and AI in 2024 Observations
Anthropic reports that election-related activity accounted for less than 1% of Claude's total usage during the 2024 election cycle, supported by proactive safety policies and the new Clio analysis tool.
LeMaterial v1.0: LeMat-Bulk dataset release
LeMaterial v1.0 launches as an open-source initiative releasing the LeMat-Bulk dataset, which unifies 6.7M entries from Materials Project, Alexandria, and OQMD into a standardized format with seven properties to accelerate materials discovery.
Sora is here – OpenAI releases video generation model Sora Turbo with new interface and subscription access
OpenAI announced the release of Sora Turbo, a faster video generation model available as a standalone product for ChatGPT Plus and Pro users, featuring a new interface, up to 1080p resolution and 20‑second clips, and built‑in safety measures.
Hugging Face Open Preference Dataset for Text-to-Image Generation
The Data is Better Together community has released an Apache 2.0 licensed open preference dataset for text-to-image generation to address the lack of open-source preference data for model alignment.
Hugging Face Models in Amazon Bedrock Marketplace
Hugging Face has integrated 83 open models into the Amazon Bedrock Marketplace, allowing AWS customers to deploy open models on managed infrastructure while maintaining compatibility with Bedrock APIs.
OpenAI Sora: Creative Workflow Integration for Animator Lyndon Barrois
Animator Lyndon Barrois uses OpenAI Sora to bypass traditional studio production pipelines, enabling the direct translation of imagination into high-fidelity video content.
Vallée Duhamel and Sora
OpenAI highlights the artistic perspective of Vallée Duhamel on the Sora video generation model, though the product is noted as no longer available as of April 26, 2026.
OpenAI Sora System Card
OpenAI has released a system card for Sora, its video generation model capable of producing videos up to 20 seconds at 1080p resolution using a diffusion-transformer architecture.
Minne Atairu and Sora
Interdisciplinary artist Minne Atairu utilizes OpenAI's Sora to challenge patriarchal imagery and redefine cultural icons.
OpenAI Product Team AI Integration
OpenAI has released a webinar and guidance on integrating AI into product team workflows to enhance productivity and product development.
Grok Image Generation Release
xAI has released Aurora, an autoregressive mixture-of-experts image generation model for Grok that supports photorealistic rendering, precise text following, and multimodal image editing.
Ollama Structured Outputs Release
Ollama now supports structured outputs, allowing users to constrain model responses to a specific format defined by a JSON schema for increased reliability and consistency.
ChatGPT Pro Release
OpenAI has launched ChatGPT Pro, a $200 monthly subscription plan providing unlimited access to o1, o1-mini, GPT-4o, and a high-compute 'o1 pro mode' for complex problem solving.
OpenAI o1 System Card
OpenAI released the o1 system card detailing its chain-of-thought reasoning model, its safety evaluations, and its Preparedness Framework ratings of medium risk for persuasion and CBRN, low for cybersecurity and model autonomy.
PaliGemma 2 Release Notes
Google has released PaliGemma 2, a new family of vision language models that combine the SigLIP image encoder with the Gemma 2 text decoder across three parameter sizes and multiple input resolutions.
How good are LLMs at fixing their mistakes? A chatbot arena experiment with Keras and TPUs
Hugging Face tested several sub‑10B LLMs on a simple calendar‑API task and found that Gemma 2 9B consistently fixed mistakes with minimal prompting, while smaller and older models struggled or required many corrective turns.
OpenAI and Future Strategic Partnership for Specialist Content
OpenAI and Future have partnered to integrate content from Future's 200-plus specialist media brands into ChatGPT, providing users with reliable, expert information and expanding the publisher's distribution reach.
Morgan Stanley AI Integration and Evaluation Framework
Morgan Stanley collaborated with OpenAI to deploy GPT-4 and Whisper powered tools, achieving 98% advisor adoption through a rigorous evaluation framework focused on reliability and compliance.
AraGen Benchmark and Leaderboard: Introducing 3C3H Evaluation for Arabic LLMs
Hugging Face introduced AraGen, a dynamic benchmark and leaderboard for Arabic LLMs that uses the 3C3H measure to evaluate correctness, completeness, conciseness, helpfulness, honesty, and harmlessness.
Hugging Face CFM Case Study: Fine-tuning Small Models with LLM Insights
Capital Fund Management (CFM) improved financial Named Entity Recognition (NER) accuracy by up to 6.4% and reduced inference costs by up to 80x by using Llama 3.1 to assist in labeling data for fine-tuning compact models like GLiNER and SpanMarker.
Open Source Developers Guide to the EU AI Act
The Hugging Face guide explains how the EU AI Act applies to open source AI developers, outlining obligations for limited‑risk AI systems and non‑systemic‑risk general purpose AI models and pointing to tools for compliance.
QwQ-32B-Preview: Exploring Deep Reasoning Capabilities
Qwen has introduced QwQ-32B-Preview, a model designed for deep reasoning and complex problem-solving through an internal chain-of-thought process.