✷ The archive · 11 labs · 3,060 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Genmab "AI Everywhere" rollout: enterprise ChatGPT expansion and 100+ custom GPTs for biotech
Genmab expanded ChatGPT Enterprise to over 2,000 staff and deployed 100+ custom GPTs, saving each user 3.5 hours weekly and accelerating biotech research and documentation.
Anthropic Contextual Retrieval technique boosts RAG accuracy
Anthropic announced Contextual Retrieval, a preprocessing method that adds chunk‑specific context to embeddings and BM25 indexes, cutting top‑20 retrieval failures by up to 67% and enabling cheaper, more reliable Retrieval‑Augmented Generation.
Anthropic Contextual Retrieval
Anthropic introduces Contextual Retrieval, a method that reduces RAG retrieval failures by up to 67% by prepending chunk-specific context to data before embedding.
Qwen2.5 Release Notes: New Foundation, Coder, and Math Models
Qwen has released Qwen2.5, a comprehensive suite of open-weight dense decoder-only models including general-purpose LLMs, specialized Coder and Math variants, and an updated Qwen2-VL-72B.
Qwen2.5 LLM series release
Qwen announced the Qwen2.5 series – open-source decoder-only LLMs from 0.5B to 72B parameters that double the capability of Qwen2 while adding a larger 18‑trillion‑token dataset, major gains in knowledge, coding, math, and alignment, and a 128K token context window.
Qwen2.5-Coder Release Notes
Qwen has released Qwen2.5-Coder, a series of open-source coding models trained on 5.5 trillion tokens that outperform larger models in code generation, reasoning, and mathematics.
Qwen2.5-Math Release Notes
Qwen has released Qwen2.5-Math, a series of open-source mathematical LLMs supporting bilingual reasoning and Tool-Integrated Reasoning (TIR) to outperform leading closed-source models on complex math benchmarks.
Fine-tuning LLMs to 1.58-bit: Extreme Quantization with BitNet
Hugging Face demonstrates that existing LLMs, such as Llama 3 8B, can be fine-tuned to 1.58-bit ternary precision using a dynamic warmup quantization strategy, significantly reducing memory and energy requirements while maintaining strong performance.
Bespoke-Minicheck: Reducing LLM Hallucinations via Grounded Factuality Checking
Ollama has integrated Bespoke-Minicheck, a grounded factuality checking model from Bespoke Labs that detects hallucinations by verifying claims against source documents.
Arco Educação and OpenAI Partner to Enhance Teaching in Brazil
Arco Educação is partnering with OpenAI to implement GPT-4 powered tools, specifically a Teacher Assistant, to reduce administrative burdens and create personalized lesson plans for students with diverse learning needs in Brazil.
Mistral AI September 2024 Release
Mistral AI has introduced a free tier for la Plateforme, significant price reductions across its model family, the Mistral Small v24.09 model, and free vision capabilities via Pixtral 12B on le Chat.
Pixtral 12B Release Notes
Mistral AI has released Pixtral 12B, a natively multimodal model that combines a new 400M parameter vision encoder with a 12B parameter decoder to deliver high-performance multimodal reasoning without compromising text capabilities.
Hugging Face SQL Console for Datasets
Hugging Face has introduced a browser-based SQL Console powered by DuckDB WASM that allows users to query, filter, and transform datasets directly on the Hub without backend dependencies.
OpenAI Safety and Security Practices Update
OpenAI has established an independent Board oversight committee to govern critical safety and security measures for model development and deployment.
HuggingChat Community Tools Release
Hugging Face has introduced Community Tools on HuggingChat, allowing users to integrate any public Hugging Face Space as a tool for LLMs to use directly within the chat interface.
Accelerate 1.0.0 Release Candidate Announcement
Accelerate 1.0.0 release candidates add FP8, DeepSpeed multi‑model, torch.compile, and new data‑loader/pipeline features while stabilizing the API for large‑scale training and inference.
OpenAI o1-preview release notes / what's new
OpenAI has released o1-preview and o1-mini, a new series of reasoning models designed to solve complex problems in science, coding, and mathematics through extended thinking time.
OpenAI Learning to Reason with LLMs
OpenAI demonstrates the reasoning capabilities of its latest models through a detailed walkthrough of a complex cipher decoding task, illustrating the step-by-step logical progression required to solve it.
OpenAI o1-mini release notes
OpenAI has released o1-mini, a cost-efficient reasoning model optimized for STEM tasks that offers significantly lower latency and cost than o1-preview while maintaining high performance in math and coding.
OpenAI o1 Contributions
OpenAI has published a comprehensive list of the internal and external contributors who developed the OpenAI o1 model, detailing the organizational roles and safety leadership involved in its creation.
Coding with OpenAI o1
OpenAI o1 enhances software development by enabling users to build more complex and consistent code through advanced reasoning capabilities.
OpenAI o1 and Quantum Physics
OpenAI o1 is a new series of AI models designed for complex reasoning in science, coding, and math by spending more time thinking before responding.
OpenAI o1 and Economics
OpenAI o1 is a new series of AI models designed to reason through complex tasks in science, coding, math, and economics by spending more time thinking before responding.
Decoding Genetics with OpenAI o1
OpenAI o1 is a new series of AI models designed for complex reasoning in science, coding, and math, providing geneticists like Catherine Brownstein with a tool to manage the vast complexity of genomic data.
Anthropic Circuits Updates August 2024
Anthropic's August 2024 Circuits updates provide preliminary research insights into multiagent system failures, worker retraining programs, and a research version of Claude's progress on the Riemann hypothesis.
Ada Customer Service Automation with GPT-4
Ada has rebuilt its AI-native customer service platform using GPT-4, doubling its resolution rate from 30% to up to 60% (and over 80% for top performers).
Hugging Face and TruffleHog Partnership for Secret Scanning
Hugging Face has partnered with Truffle Security to integrate TruffleHog's secret scanning capabilities into its automated pipeline and provide a native scanner for users to proactively scan their own account data.
Salesforce Integrates Anthropic Claude AI
Salesforce has integrated Anthropic's Claude 3.5 Sonnet, Claude 3 Opus, and Claude 3 Haiku models into its platform via Amazon Bedrock, allowing enterprises to customize AI-powered CRM applications.
Qwen2-VL release: open-source 2B/7B vision-language models and 72B API with state-of-the-art image, video, and multilingual capabilities
Qwen released Qwen2-VL, a new vision-language model series (2B, 7B open-source, 72B API) that sets state-of-the-art performance on image, document, multilingual, and long-video understanding while supporting visual agent capabilities.
LeRobotDataset video‑encoding format reduces robotics dataset size and speeds up training
Hugging Face released the LeRobotDataset video‑encoding format, shrinking robotics visual data to about 14 % of its original size while keeping loading speed and training performance intact.
Arizona State University Integrates ChatGPT Edu for Personalized Education
Arizona State University has deployed ChatGPT Edu across more than 200 projects to personalize learning, advance research, and enhance operational efficiency for its 181,000 students.
Hugging Face blog post highlights five under‑rated Hub tools and a free semantic‑search use case
Hugging Face announced five under‑rated Hub tools—ZeroGPU, multi‑process Docker, Gradio API, webhooks, and Nomic Atlas—and showed how to combine them into a free, auto‑updating semantic‑search app for Reddit data.
Hugging Face Training Efficiency: Packing with Flash Attention 2
Hugging Face has introduced boundary-aware packing for instruction tuning examples, enabling up to 2x training throughput increase and 20% peak memory reduction when used with Flash Attention 2.
OpenAI and Condé Nast Partnership for Content Integration
OpenAI has partnered with Condé Nast to integrate content from brands like Vogue and The New Yorker into ChatGPT and the SearchGPT prototype to improve information discovery and source attribution.
Upwork OpenAI Integration and AI-First Strategy
Upwork has transitioned to an OpenAI-centric ecosystem, integrating GPT-4o and GPT-3.5 into its marketplace features, internal fraud detection, and corporate productivity via ChatGPT Enterprise.
OpenAI launches fine‑tuning for GPT‑4o with free token quota and announces platform sunset
OpenAI launched fine‑tuning for GPT‑4o, enabling developers to customize the model with minimal data and offering free training tokens until September 23, while announcing the platform will close to new users after May 8 2026.
Deploying Meta Llama 3.1 405B on Google Cloud Vertex AI
Hugging Face provides a guide for programmatically deploying the FP8 quantized version of Meta Llama 3.1 405B on Google Cloud Vertex AI using Text Generation Inference (TGI) and A3 machine series.
OpenAI Disrupts Iranian Influence Operation Storm-2035
OpenAI has banned a cluster of ChatGPT accounts used by the Iranian influence operation Storm-2035 to generate political content for the 2024 U.S. election and other global events.
Indeed Contextual Job Matching with OpenAI
Indeed integrated OpenAI's GPT models to provide personalized explanations for job recommendations, resulting in a 20% increase in started job applications and a 13% uplift in downstream success.
OpenAI and The Met Museum Collaboration: Awakening Sleeping Beauties
OpenAI collaborated with the Metropolitan Museum of Art's Costume Institute to create an AI-powered chat experience allowing visitors to interact with a historical figure, Natalie Potter, based on curated historical datasets.
Hugging Face Infini-Attention Reproduction Analysis
Hugging Face's attempt to reproduce Infini-Attention found that while gating convergence can be improved, the method's performance degrades with increased memory compression and remains less reliable than Ring Attention, YaRN, or RoPE scaling.
OpenAI SWE-bench Verified release: human‑validated benchmark improves software‑engineering evaluation
OpenAI released SWE‑bench Verified, a human‑validated 500‑sample subset of the SWE‑bench software‑engineering benchmark that removes ambiguous issues and unfair tests, enabling more reliable evaluation of AI models’ coding abilities; GPT‑4o solves 33.2% of these samples, more than double its score on the original benchmark.
Introduction to ggml
ggml is a lightweight, C/C++ machine learning library optimized for Transformer inference and on-device LLM execution across diverse hardware backends.
Grok-2 Beta Release
xAI has released Grok-2 and Grok-2 mini in beta, featuring frontier-level capabilities in chat, coding, and reasoning that outperform Claude 3.5 Sonnet and GPT-4 Turbo on the LMSYS leaderboard.
Hugging Face Unified Tool Use API
Hugging Face has introduced a unified tool use API that allows developers to use the same code to implement tool calling across Mistral, Cohere, NousResearch, and Llama models.
Falcon Mamba 7B Release Notes
The Technology Innovation Institute (TII) has released Falcon Mamba 7B, the first large-scale pure State Space Language Model (SSLM) that matches the performance of state-of-the-art transformer models while eliminating attention-based memory scaling issues.
Qwen2-Audio Release Notes
Qwen2-Audio is an audio-language model capable of voice chat and audio analysis across more than eight languages and dialects, surpassing previous state-of-the-art performance on multiple benchmarks.
Zico Kolter Joins OpenAI Board of Directors
OpenAI has appointed Zico Kolter, a professor and Director of the Machine Learning Department at Carnegie Mellon University, to its Board of Directors and Safety and Security Committee.
Hugging Face acquires XetHub to upgrade Hub storage and collaboration
Hugging Face acquired XetHub to replace Git LFS with a more efficient storage backend, enabling incremental updates, trillion‑parameter model support, and better collaboration on massive AI datasets.
GPT‑4o System Card – capabilities, safety mitigations, and risk assessment
OpenAI released the GPT‑4o System Card, detailing the omni‑model's capabilities, safety mitigations, and medium overall risk rating.