✷ The archive · 11 labs · 3,060 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Anthropic expands model safety bug bounty program
Anthropic announced an expanded, invite‑only bug bounty program that rewards up to $15,000 for discovering universal jailbreak attacks on its next‑generation AI safety mitigations, aiming to protect high‑risk domains such as CBRN and cybersecurity.
Qwen2-Math Release Notes
Qwen has released Qwen2-Math, a series of specialized mathematical LLMs (1.5B, 7B, and 72B) that outperform several closed-source models, including GPT-4o, on math benchmarks.
Rakuten and OpenAI Partnership: Leveraging Generative AI for Customer Insights
Rakuten is utilizing OpenAI's APIs, RAG, and Code Interpreter to transform unstructured data into automated customer service, review summaries, and B2B market insights.
Mistral AI Model Customization and Agents Release
Mistral AI has introduced model customization for flagship models including Mistral Large 2 and Codestral, alongside an alpha release of Agents and the stable 1.0 version of their client SDK.
OpenAI Structured Outputs API feature announcement
OpenAI introduced Structured Outputs on August 6, 2024, a new API feature that guarantees model responses exactly match developer‑provided JSON Schemas, improving reliability for data‑centric applications.
Hugging Face TextImage Augmentation pipeline release
Hugging Face and Albumentations AI released a TextImage Augmentation pipeline that jointly modifies document images and their text, enabling realistic synthetic data generation and robust fine‑tuning of vision‑language models on limited document datasets.
Hugging Face 2024 Security Feature Highlights
Hugging Face has detailed its 2024 security landscape, introducing a suite of default protections for all users and advanced governance controls for Enterprise Hub users.
Claude AI Availability in Brazil
Anthropic has expanded the availability of Claude, its AI assistant, to consumers and businesses in Brazil via web, mobile apps, and API.
Google releases Gemma 2 2B, ShieldGemma, and Gemma Scope
Google released Gemma 2 2B, ShieldGemma safety classifiers, and Gemma Scope sparse autoencoders, expanding open‑source LLM capabilities, moderation tools, and interpretability resources.
Anthropic Circuits Updates July 2024
Anthropic released a series of preliminary research updates in July 2024 focusing on interpretability, multiagent system failures, and mathematical breakthroughs in an unreleased Claude research version.
Quanto quantization cuts memory for Transformer diffusion pipelines
Hugging Face quantization (Quanto) reduces GPU memory for Transformer diffusion models from ~12 GB to ~5 GB with minimal latency and quality impact.
Hugging Face NVIDIA NIM API (serverless) launch and deprecation
Hugging Face launched the NVIDIA NIM API (serverless) for Enterprise Hub users, enabling pay‑as‑you‑go, serverless inference of open‑source LLMs on NVIDIA DGX Cloud H100 GPUs via an OpenAI‑compatible API.
LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs
Hugging Face introduces LAVE (LLM-Assisted VQA Evaluation) to address the rigidity of traditional VQA metrics, demonstrating a 50% accuracy gain in evaluating zero-shot performance on the Docmatix dataset.
OpenAI SearchGPT Prototype Announcement
OpenAI has introduced SearchGPT, a temporary prototype that combines AI models with real-time web information to provide direct answers with clear source citations.
Ollama Tool Support Release
Ollama has introduced tool calling support for models like Llama 3.1, allowing local LLMs to interact with external APIs, functions, and code interpreters.
Mistral Large 2 release notes / what's new
Mistral AI has released Mistral Large 2, a 123B parameter model designed for high throughput single-node inference with competitive performance in coding, reasoning, and multilingual capabilities.
OpenAI Rule-Based Rewards for Model Safety
OpenAI has introduced Rule-Based Rewards (RBRs), a method that uses explicit, step-by-step rules to align AI safety behavior without requiring extensive human data collection.
Llama 3.1 Release Notes: Multilinguality, Long Context, and 405B Model
Meta has released Llama 3.1, featuring models in 8B, 70B, and 405B sizes with 128K context length, multilingual support for 8 languages, and a permissive license allowing synthetic data generation.
Running Mistral 7B with Core ML
Hugging Face demonstrates how to run Mistral 7B on Mac using new Core ML features from WWDC 24, achieving a model size reduction to under 4GB using 4-bit block-wise quantization.
GPT-4o mini release notes / what's new
OpenAI has released GPT-4o mini, a highly cost-efficient small model that outperforms GPT-3.5 Turbo and other small models on key reasoning and multimodal benchmarks.
Mistral NeMo Release Notes
Mistral AI and NVIDIA have released Mistral NeMo, a 12B parameter model featuring a 128k context window and a new, highly efficient multilingual tokenizer called Tekken.
Docmatix Dataset Release
Hugging Face has released Docmatix, a Document Visual Question Answering (DocVQA) dataset featuring 2.4 million images and 9.5 million Q/A pairs, providing a 240x increase in scale over previous datasets.
TGI Multi-LoRA: Deploy Once, Serve 30 Models
Hugging Face introduces Multi-LoRA serving in Text Generation Inference (TGI), allowing organizations to deploy a single base model and dynamically serve dozens of specialized fine-tuned adapters to reduce cost and operational complexity.
ChatGPT Enterprise Compliance and Administrative Tools Update
OpenAI has introduced a Compliance API, SCIM for automated user management, and granular GPT controls to help enterprise customers meet regulatory requirements and scale AI deployments securely.
OpenAI Prover‑Verifier Games improve legibility of language model outputs
OpenAI announced a prover‑verifier game training method that makes strong language models generate solutions that weaker models can verify, improving both correctness and human legibility.
Anthropic and Menlo Ventures Launch $100 Million Anthology Fund
Anthropic and Menlo Ventures have launched the $100 million Anthology Fund to accelerate the development of AI applications across infrastructure, industry-specific tools, and trust and safety.
Codestral Mamba Release Notes
Mistral AI has released Codestral Mamba, a 7B parameter model utilizing the Mamba architecture to provide linear time inference and high-performance coding and reasoning capabilities.
Mathstral 7B Release
Mistral AI has released Mathstral, a 7B parameter model specializing in STEM subjects and advanced mathematical reasoning, developed in collaboration with Project Numina.
Argilla SDK Chatbot with distilabel – End‑to‑End Tutorial
Hugging Face released a tutorial showing how to build an Argilla 2.0 chatbot using distilabel‑generated synthetic data, fine‑tuned embeddings, lancedb vector storage, and a Gradio app deployed on Spaces.
SmolLM Release: High-Performance Small Language Models
Hugging Face introduces SmolLM, a family of state-of-the-art small language models (135M, 360M, and 1.7B parameters) trained on a meticulously curated high-quality dataset.
NuminaMath 7B TIR wins AIMO Progress Prize – technical recap
NuminaMath 7B TIR won the first AIMO Progress Prize by solving 29 of 50 hidden math problems, showcasing the power of a two-stage fine‑tuning recipe, large high‑quality math data, and a self‑consistency with tool‑integrated reasoning inference strategy.
OpenAI and Los Alamos National Laboratory Bioscience Research Partnership
OpenAI and Los Alamos National Laboratory have partnered to evaluate how multimodal frontier AI models like GPT-4o can safely assist scientists in physical laboratory settings to advance bioscientific research.
Hugging Face PII Detection Experiment with Presidio
Hugging Face is experimenting with integrating Microsoft Presidio into the Dataset Hub to provide automatic PII detection reports, helping practitioners identify and mitigate privacy risks in ML datasets.
TRL adds Direct Preference Optimization support for Vision‑Language Models
Hugging Face added Direct Preference Optimization (DPO) support for Vision‑Language Models in the TRL library, enabling fine‑tuning of models like Idefics‑2 with preference data using bfloat16 quantization and LoRA to fit on a single GPU.
Hugging Face and KerasHub Integration
Hugging Face and KerasHub now share a model save format, allowing KerasHub users to directly load over 300,000 Transformers library models from the Hugging Face Hub.
Google Cloud TPUs on Hugging Face Inference Endpoints and Spaces
Hugging Face has integrated Google Cloud TPU v5e support into Inference Endpoints and Spaces, enabling users to deploy and scale AI models with cost-effective, high-performance hardware.
Banque des Territoires, Polyconseil, and Hugging Face Deploy Sovereign RAG Solution for EduRénov Program
Hugging Face, Polyconseil, and Banque des Territoires launched a sovereign, open‑source RAG system to automate email support for France's EduRénov school‑renovation program, ensuring data residency while scaling generative AI for public policy.
Hugging Face Dataset Hub Search Features Update
Hugging Face has introduced four new search filters—modality, size, format, and library compatibility—to improve the discoverability of over 180,000 public datasets on the Dataset Hub.
Accelerating Protein Language Model ProtST on Intel Gaudi 2
Intel and MILA have optimized the ProtST multi-modal protein language model for Intel Gaudi 2 accelerators, achieving up to 2.92x faster fine-tuning compared to NVIDIA A100 GPUs.
Hugging Face Transformers Code Agent GAIA Benchmark Results
Hugging Face achieved a top ranking on the GAIA benchmark using a Code Agent built with the transformers.agents library, demonstrating that code-based actions are more efficient and effective than JSON-based tool calling.
Anthropic Third-Party Model Evaluation Initiative
Anthropic has launched a funding initiative to support third-party organizations in developing high-quality evaluations to measure advanced AI capabilities and safety risks.
Anthropic Circuits Updates June 2024
Anthropic's June 2024 Circuits updates provide preliminary research findings on interpretability, multiagent system failures, and advancements in mathematical capabilities regarding the Riemann zeta function.
Finding GPT-4’s mistakes with GPT-4
OpenAI introduced CriticGPT, a GPT-4 based model designed to identify errors in ChatGPT's code output to assist human trainers during the RLHF process.
OpenAI and TIME Strategic Content Partnership
OpenAI and TIME have entered a multi-year strategic partnership to integrate TIME's 101-year archive of journalism into OpenAI products and provide TIME with access to OpenAI technology for product development.
Google Gemma 2 Release Notes
Google has released Gemma 2, a family of open-weight LLMs available in 9B and 27B parameter sizes, featuring technical advances in distillation and attention mechanisms to improve performance.
Google Gemma 2 release on Ollama
Google Gemma 2, available in 2B, 9B, and 27B parameter variants, launches on Ollama with a new architecture that delivers class‑leading performance and efficiency, outperforming larger open models.
Anthropic Expands Claude Access for Government Agencies
Anthropic has made Claude 3 Haiku and Claude 3 Sonnet available to the US Intelligence Community and AWS GovCloud to support government missions through secure, specialized service agreements.
XLSCOUT ParaEmbed 2.0 Release
XLSCOUT has released ParaEmbed 2.0, a proprietary embedding model fine-tuned on expert-curated patent data that achieves a 23% increase in accuracy over ParaEmbed 1.0.
Anthropic Introduces Claude Projects for Collaborative AI Workflows
Anthropic has launched Projects for Claude Pro and Team users, allowing them to organize chats with curated knowledge sets and custom instructions within a 200K context window.
Hugging Face Ethics and Society Newsletter #6: The Importance of Data Quality
Hugging Face outlines a holistic, responsible approach to data quality, emphasizing that high-quality data must be fit for its intended purpose to ensure AI model performance, fairness, and scientific reproducibility.