✷ The archive · 11 labs · 3,060 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Hugging Face Hub Storage Rearchitecture
Hugging Face is redesigning its upload and download architecture by introducing a content-addressed store (CAS) to enable byte-level deduplication and improve global transfer speeds for massive AI models.
SmolVLM release notes / what's new
Hugging Face introduces SmolVLM, a 2B parameter Vision Language Model (VLM) that is fully open-source and optimized for low memory footprints and high throughput on edge devices.
You could have designed state of the art positional encoding
The Hugging Face blog post walks through an iterative design of positional encoding for transformers, showing how sinusoidal encoding leads to Rotary Positional Encoding (RoPE) and why it matters for modeling token relationships.
Anthropic Model Context Protocol (MCP) Release
Anthropic has open-sourced the Model Context Protocol (MCP), a universal standard for connecting AI assistants to data sources like content repositories, business tools, and development environments to eliminate fragmented integrations.
Ollama Python Library 0.4 Release Notes
Ollama Python library 0.4 introduces the ability to pass Python functions directly as tools, automating JSON schema generation via Pydantic and docstring parsing.
Anthropic and AWS Expand Partnership for AI Development
Anthropic and AWS have expanded their collaboration, featuring a new $4 billion investment from Amazon and a deep technical partnership to optimize Trainium hardware for future foundation models.
Advancing red teaming with people and AI – OpenAI's approach and application to the o1 family
OpenAI published a white paper detailing its external red teaming process for AI models and applied it to prepare the OpenAI o1 family for public release.
Grab GPT-4o Vision Fine-Tuning for GrabMaps
Grab has implemented GPT-4o vision fine-tuning to automate the localization of traffic signs and lane counting for GrabMaps, improving speed limit sign localization by 13% and lane count accuracy by 20%.
Hugging Face and LLM-jp Launch Open Japanese LLM Leaderboard
Hugging Face and LLM-jp have introduced the Open Japanese LLM Leaderboard, a transparent evaluation platform featuring over 20 datasets to benchmark the performance of Japanese large language models.
Hugging Face Introduces Content-Defined Chunking to Improve Storage Efficiency
Hugging Face announced a content-defined chunking storage approach via its Xet team that reduces storage and transfer costs for large model and dataset files by only uploading modified chunks.
Faster Text Generation with Self-Speculative Decoding
Hugging Face introduces self-speculative decoding via LayerSkip, a method that uses a single LLM's early layers for drafting and later layers for verification to increase generation speed and reduce memory overhead.
FlagEval Debate: A New Multilingual LLM Evaluation Framework
BAAI has launched FlagEval Debate, a dynamic evaluation platform where LLMs compete in multilingual debates to better assess reasoning, logic, and adversarial capabilities.
Rox Revenue Platform Integration with OpenAI
Rox has launched an AI-powered revenue management platform using OpenAI's API to automate data unification, sales workflows, and account monitoring through a system of AI agent swarms.
Hugging Face Judge Arena: Benchmarking LLMs as Evaluators
Hugging Face has launched Judge Arena, a crowdsourced platform that uses human voting to determine which LLMs are most effective as evaluators for grading other AI-generated responses.
Anthropic A Statistical Approach to Model Evaluations
Anthropic proposes a rigorous statistical framework for AI model evaluations to distinguish real capability differences from random noise using tools like the Central Limit Theorem and power analysis.
Mistral AI le Chat Updates November 2024
Mistral AI has introduced several beta features to le Chat, including web search with citations, a collaborative Canvas interface, and multimodal capabilities powered by Pixtral Large.
Pixtral Large Release Notes
Mistral AI has released Pixtral Large, a 124B open-weights multimodal model that achieves state-of-the-art performance on MathVista, DocVQA, and VQAv2 while maintaining the text capabilities of Mistral Large 2.
OpenAI Opens Paris Office to Expand French AI Ecosystem
OpenAI has opened a new office in Paris to support the rapid adoption of AI across French organizations, startups, and government collaborations.
Qwen2.5-Turbo 1M Token Context Length Release
Qwen has released Qwen2.5-Turbo, which extends the model's context window to 1 million tokens while significantly improving inference speed and maintaining competitive performance on short-sequence tasks.
The Estée Lauder Companies ChatGPT Enterprise Implementation
The Estée Lauder Companies has deployed ChatGPT Enterprise to analyze 75+ years of consumer and clinical data, creating over 240 custom GPTs to accelerate product development and market responsiveness.
Hugging Face Hub Dataset Sharing for Researchers
Hugging Face Hub provides a comprehensive platform for hosting and sharing large-scale ML datasets with integrated tools for exploration, security, and community engagement.
Qwen2.5-Coder Series Release Notes
Qwen has released the Qwen2.5-Coder series, featuring six model sizes from 0.5B to 32B, with the 32B-Instruct model achieving SOTA open-source performance comparable to GPT-4o in coding tasks.
Mistral Batch API Release
Mistral AI has launched a Batch API that allows users to process high-volume requests at a 50% lower cost than synchronous API calls.
Mistral Moderation API Release
Mistral AI has released a new content moderation API, an LLM-based classifier that categorizes text into nine safety categories across multiple languages to provide system-level guardrails for AI deployments.
Llama 3.2 Vision available in Ollama
Ollama has released support for Llama 3.2 Vision in 11B and 90B parameter sizes, enabling local execution of multimodal capabilities including OCR and image analysis.
Hugging Face PyCharm Integration
Hugging Face has integrated its Hub directly into PyCharm Professional, allowing developers to discover, insert, and manage machine learning models without leaving their IDE.
Argilla 2.4 release notes / what's new
Argilla 2.4 introduces a no-code UI for importing Hugging Face Hub datasets to build fine-tuning and evaluation datasets through human feedback.
xAI API Public Beta Release
xAI has launched a public beta of its API, providing developers programmatic access to the Grok series of foundation models, starting with the grok-beta model.
Claude 3 Haiku Fine-Tuning in Amazon Bedrock
Anthropic has made fine-tuning for Claude 3 Haiku generally available in Amazon Bedrock, allowing users to customize the model for specialized tasks to achieve higher accuracy and lower costs.
Introducing ChatGPT search
OpenAI has integrated a new web search capability into ChatGPT, allowing users to get fast, timely answers with direct links to relevant web sources.
Promega ChatGPT Adoption Case Study
Promega has integrated ChatGPT across its organization, deploying over 1,400 custom GPTs to accelerate manufacturing, sales, and marketing workflows.
Anthropic: The Case for Targeted AI Regulation
Anthropic advocates for urgent, narrowly-targeted government regulation of frontier AI models within the next 18 months to mitigate catastrophic cyber and CBRN risks while preserving innovation.
OpenAI SimpleQA benchmark release
OpenAI released SimpleQA, an open‑source benchmark of 4,326 short fact‑seeking questions for evaluating the factual accuracy and calibration of frontier language models.
Decagon Customer Support Automation with OpenAI
Decagon utilizes a multi-model strategy featuring GPT-3.5, GPT-4, and o1-mini to automate up to 91% of global support for enterprise clients.
Universal Assisted Generation: Faster Decoding with Any Assistant Model
Hugging Face and Intel Labs introduced Universal Assisted Generation (UAG), a method that accelerates LLM inference by 1.5x-2.0x by allowing any small model to act as an assistant regardless of its tokenizer.
Claude 3.5 Sonnet Integration with GitHub Copilot
Anthropic's Claude 3.5 Sonnet is now available in public preview on GitHub Copilot, providing developers with a high-performance coding model that outperforms other publicly available models on SWE-bench Verified.
Digital Green Farmer.chat: Bolstering RAG with LLM-as-a-Judge
Digital Green implemented an LLM-as-a-judge evaluation framework for Farmer.chat, a RAG-based agricultural chatbot, to objectively measure RAG accuracy and optimize model selection across 340k queries.
Anthropic Evaluating Feature Steering: A Case Study in Mitigating Social Biases
Anthropic researchers found that while feature steering can target specific social biases and political stances in Claude 3 Sonnet, it often produces unpredictable off-target effects and degrades model capabilities outside a specific steering range.
Aya Expanse Release: Advancing Multilingual LLM Performance
Hugging Face and Cohere For AI have released Aya Expanse, a family of 8B and 32B open-weight models that set new state-of-the-art benchmarks for multilingual performance.
Simplifying, stabilizing, and scaling continuous-time consistency models
OpenAI introduces sCM, a simplified continuous-time consistency model that achieves diffusion-level sample quality in just two sampling steps, providing a ~50x speedup in generation.
HUGS launch: zero‑configuration, hardware‑optimized inference for open LLMs
Hugging Face launched HUGS, a zero‑configuration, hardware‑optimized inference service for open‑source LLMs that runs on NVIDIA, AMD, and soon AWS Inferentia and Google TPUs, enabling enterprises to host models in‑house with an OpenAI‑compatible API.
CinePile 2.0 release: adversarial refinement boosts video QA dataset quality
CinePile 2.0 introduces an adversarial refinement pipeline that upgrades weak QA pairs into vision‑dependent questions, releasing both the improved dataset and the full code, and shows significant performance gains for commercial and open‑source video‑LLMs.
SynthID Text Integration in Transformers v4.46.0
Google DeepMind and Hugging Face have integrated SynthID Text into Transformers v4.46.0, providing a method to apply imperceptible watermarks to AI-generated text for detection via trained classifiers.
OpenAI and Microsoft Partner with Lenfest Institute for AI Collaborative and Fellowship Program
OpenAI and Microsoft have partnered with the Lenfest Institute for Journalism to provide $10 million in funding and credits to help local newsrooms implement AI for business sustainability and innovation.
Deploying Speech-to-Speech on Hugging Face Inference Endpoints
Hugging Face provides a guide for deploying its Speech-to-Speech (S2S) pipeline using custom Docker images on Inference Endpoints to handle high computational demands and reduce latency.
Outlines-core 0.1.0 release notes / what's new
Hugging Face and dottxt have released outlines-core 0.1.0, a Rust port of the Outlines core algorithms for structured generation that improves index compilation speed and portability.
Stable Diffusion 3.5 Large Integration with Diffusers
Hugging Face has integrated Stable Diffusion 3.5 Large, an 8B parameter model available in standard and timestep-distilled versions, into the Diffusers library.
Hugging Face partners with Protect AI to add Guardian scanner for model security
Hugging Face partnered with Protect AI to embed the Guardian scanner into the Hub, automatically detecting dangerous model serialization exploits and improving security for the entire ML community.
Transformers.js v3 release adds WebGPU acceleration, expanded model support, and server‑side JavaScript compatibility
Transformers.js v3 adds WebGPU acceleration, new quantization formats, support for 120 model architectures, and Node.js/Deno/Bun compatibility, enabling fast, on‑device inference in browsers and JavaScript runtimes.
Anthropic Claude 3.5 Sonnet and Claude 3.5 Haiku Release
Anthropic has released an upgraded Claude 3.5 Sonnet and the new Claude 3.5 Haiku, alongside a public beta for a groundbreaking 'computer use' capability that allows the model to interact with standard computer interfaces.