NPHardEval Leaderboard: Evaluating LLM Reasoning via Computational Complexity
Hugging Face introduces the NPHardEval leaderboard, a dynamic benchmark that uses computational complexity classes to quantitatively measure the logical reasoning abilities of Large Language Models.
PatchTST Integration in Hugging Face
Hugging Face has integrated PatchTST, a Transformer-based model that uses time series patching and channel-independence to improve long-term forecasting and enable transfer learning.
Hugging Face Text Generation Inference now supports AWS Inferentia2
Hugging Face announced the general availability of Text Generation Inference on AWS Inferentia2 via Amazon SageMaker, enabling cost‑effective, high‑throughput LLM serving as an alternative to GPU deployments.
Constitutional AI with Open LLMs
Hugging Face introduces an end-to-end recipe and the llm-swarm tool to implement Constitutional AI (CAI) on open models, enabling scalable self-alignment based on user-defined principles without expensive human feedback.
OpenAI Biological Threat Evaluation: Building an Early Warning System
OpenAI conducted a human-centric study to determine if GPT-4 increases the ability of users to create biological threats compared to internet-only access, finding mild but not statistically significant uplifts in accuracy and completeness.
Enterprise Scenarios Leaderboard: Evaluating LLMs for Real-World Use Cases
Hugging Face and Patronus AI have launched the Enterprise Scenarios Leaderboard to evaluate language models on six real-world business tasks, moving beyond academic benchmarks to measure practical enterprise utility.
Accelerating StarCoder on Intel Xeon with Optimum Intel
Hugging Face and Intel demonstrate over 7x inference acceleration for the StarCoder-15B model on 4th Gen Intel Xeon processors by combining 8-bit quantization and assisted generation.
Hugging Face Hallucinations Leaderboard launch and initial findings
Hugging Face launched the Hallucinations Leaderboard to benchmark LLMs on factuality and faithfulness errors across multiple open-source datasets, offering transparent rankings that guide model selection and research.
AI Secure LLM Safety Leaderboard
Hugging Face and the Secure Learning Lab have released the LLM Safety Leaderboard, powered by the DecodingTrust framework to evaluate LLM trustworthiness across eight critical safety dimensions.
OpenAI New Embedding Models and API Updates January 2024
OpenAI has released new text-embedding-3-small and text-embedding-3-large models with improved performance and lower pricing, alongside updates to GPT-4 Turbo and GPT-3.5 Turbo.
Qwen-VL-Plus and Qwen-VL-Max Release
Qwen has released Qwen-VL-Plus and Qwen-VL-Max, large visual language models that match GPT-4V and Gemini Ultra in multimodal tasks and outperform them in Chinese text comprehension.
Hugging Face and Google Cloud Strategic Partnership
Hugging Face and Google Cloud have entered a strategic partnership to democratize machine learning by integrating open models with Google Cloud's AI infrastructure and hardware.
Open-source LLMs as LangChain Agents
Hugging Face demonstrates that open-source LLMs, specifically Mixtral-8x7B, are now capable of powering agent workflows and can outperform GPT-3.5 in general-purpose reasoning tasks.
Introducing Qwen: A Comprehensive LLM and LMM Framework
Qwen is a project towards AGI consisting of a series of multilingual large language models (LLMs) and large multimodal models (LMMs), including open-source versions ranging from 1.8B to 72B parameters.
Fine-Tuning Wav2Vec2-BERT for Low-Resource ASR
Hugging Face demonstrates how to fine-tune Meta's Wav2Vec2-BERT model for Automatic Speech Recognition (ASR) in low-resource languages, achieving performance comparable to Whisper-large-v3 while being significantly faster and more resource-efficient.
PatchTSMixer added to Hugging Face Transformers – release and quick‑start guide
PatchTSMixer, a lightweight MLP‑Mixer time‑series model from IBM Research, is now released in Hugging Face Transformers, offering state‑of‑the‑art forecasting with far lower memory and runtime costs.
Preference Tuning LLMs with Direct Preference Optimization Methods – Empirical Comparison of DPO, IPO, and KTO
Hugging Face evaluated DPO, IPO and KTO alignment methods on two 7B chat models, showing DPO consistently outperforms the others when the beta hyper‑parameter is properly tuned.
OpenAI Democratic Inputs to AI Grant Program Update
OpenAI has summarized the results of its Democratic Inputs to AI grant program, detailing ten innovative projects and establishing a new Collective Alignment team to integrate public inputs into model behavior.
OpenAI 2024 Election Integrity Initiatives
OpenAI implemented a multi-layered safety strategy for 2024 worldwide elections, focusing on elevating authoritative voting information, preventing deepfakes of political figures, and disrupting covert influence operations.
Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive
Hugging Face and Microsoft introduce optimizations using ONNX Runtime and Olive to achieve throughput gains up to 229% for SDXL Turbo and 120% for SD Turbo compared to PyTorch.
Run ComfyUI Workflows on Hugging Face Spaces with Gradio
Hugging Face provides a guide to converting complex ComfyUI workflows into Gradio applications for free, serverless deployment on Hugging Face Spaces ZeroGPU.
OpenAI and Digital Green Launch Farmer.Chat for Agricultural Extension
Digital Green has partnered with OpenAI to create Farmer.Chat, a generative AI tool that reduces the cost of agricultural extension services from $35 to $0.35 per farmer while supporting multiple languages in India and Kenya.
Hugging Face Leaderboard Templates: Implementing the Vectara HHEM Leaderboard
Hugging Face has released open-source leaderboard templates that enable developers to build dynamic LLM evaluation boards, as demonstrated by Vectara's new Hughes Hallucination Evaluation Model (HHEM) leaderboard.
ChatGPT Team Release Notes
OpenAI has launched ChatGPT Team, a self-serve plan designed for small to medium teams providing collaborative workspaces, admin tools, and enterprise-grade data privacy.
OpenAI GPT Store Launch
OpenAI has launched the GPT Store, allowing ChatGPT Plus, Team, and Enterprise users to discover, share, and monetize custom GPT versions of ChatGPT.
Unsloth and Hugging Face TRL Integration for Faster LLM Fine-tuning
Unsloth is a lightweight library that accelerates LLM fine-tuning by up to 2.7x and reduces memory usage by up to 74% with 0% accuracy degradation compared to QLoRA.
OpenAI and Journalism: Response to The New York Times Lawsuit
OpenAI asserts that training AI models on public internet data is fair use and describes the regurgitation of copyrighted content as a rare bug, responding to a lawsuit from The New York Times.
WHOOP Coach: Personalized Health Coaching via GPT-4
WHOOP has integrated OpenAI's GPT-4 to launch WHOOP Coach, an AI-powered fitness and health coach that provides personalized, on-demand guidance based on a user's unique physiological data.
aMUSEd: Efficient Text-to-Image Generation
Hugging Face has released aMUSEd, an efficient non-diffusion text-to-image model based on Masked Image Modeling (MIM) and an open reproduction of Google's MUSE.
Hugging Face SDXL Dreambooth LoRA Advanced Training Guide
Hugging Face introduces an advanced training script for SDXL Dreambooth LoRAs, combining Pivotal Tuning and the Prodigy optimizer to improve concept capture and image quality.
Speculative Decoding Enables 2× Faster Whisper Inference
Hugging Face demonstrates that speculative decoding halves Whisper transcription latency while preserving identical outputs and accuracy.
2023 Year of Open LLMs Review
Hugging Face’s 2023 recap shows a surge of open‑source LLM releases, smaller high‑performing models, and new fine‑tuning techniques that dramatically broaden access and community participation.
OpenAI Practices for Governing Agentic AI Systems
OpenAI proposes a framework of baseline responsibilities and safety best practices to ensure the responsible integration of agentic AI systems—AI that pursues complex goals with limited supervision—into society.
OpenAI Superalignment Fast Grants
OpenAI has launched a $10 million grants program to fund technical research into the alignment and safety of superhuman AI systems.
Summer Health uses GPT-4 to automate pediatric visit notes
Summer Health has integrated GPT-4 to transform pediatrician observations into clear, jargon-free visit notes, reducing administrative time per note from 10 minutes to 2 minutes.
OpenAI Weak-to-Strong Generalization Research
OpenAI researchers have demonstrated that a smaller, less capable model can supervise a larger model to elicit capabilities near GPT-3.5 levels, providing a potential path for aligning superhuman AI systems.
OpenAI and Axel Springer Partnership for AI Journalism
OpenAI and Axel Springer have partnered to integrate authoritative news content from brands like POLITICO and Business Insider into ChatGPT, while utilizing Axel Springer content for LLM training.
Mixture of Experts (MoE) Explained
Mixture of Experts (MoE) allows transformer models to scale parameters while maintaining efficient pretraining and faster inference by activating only a subset of neural network experts per token.
Mixtral 8x7B Release Notes
Mistral AI has released Mixtral 8x7B, a Mixture of Experts (MoE) model that outperforms Llama 2 70B and matches GPT-3.5 performance on most benchmarks while remaining commercially permissive under Apache 2.0.
SetFitABSA: Few-Shot Aspect Based Sentiment Analysis
Hugging Face and Intel Labs introduced SetFitABSA, a prompt-less, few-shot framework for Aspect-Based Sentiment Analysis that outperforms larger generative models like Llama 2 and T5 in low-data scenarios.
Optimum-NVIDIA Release Notes
Hugging Face has released Optimum-NVIDIA, an inference library that accelerates LLM inference on NVIDIA platforms by up to 28x using FP8 quantization and TensorRT-LLM.
Hugging Face LoRA dynamic loading speeds inference 300% and cuts latency
Hugging Face announced a dynamic LoRA loading system that reduces warm‑up time from 25 s to 3 s, delivering up to 300 % faster LoRA inference and cutting total response time from 35 s to 13 s.
Hugging Face and AMD GPU Acceleration for LLMs
Hugging Face and AMD have integrated out-of-the-box support for AMD Instinct GPUs into the Transformers library and Text Generation Inference, enabling high-performance LLM execution without code changes.
Hugging Face Open LLM Leaderboard DROP Benchmark Analysis
Hugging Face has removed the DROP benchmark from the Open LLM Leaderboard after discovering that flawed normalization and stop-token configurations caused most models to score incorrectly low.
OpenAI Leadership Update: Sam Altman Returns as CEO
Sam Altman has returned as CEO of OpenAI, supported by a new initial board and a commitment to strengthening corporate governance.
OpenAI Leadership Transition: Mira Murati Appointed Interim CEO
OpenAI has appointed Mira Murati as interim CEO following the departure of Sam Altman, who has left both the CEO role and the board of directors.
OpenAI Data Partnerships
OpenAI has launched Data Partnerships to collaborate with organizations to create public and private datasets for training AI models to improve domain-specific understanding and move toward AGI.
SDXL and Stable Diffusion Fast Inference with Latent Consistency LoRAs
Hugging Face introduces LCM LoRAs, a method to enable high-quality image generation in 4 to 8 steps for SDXL and Stable Diffusion models, significantly reducing inference time.
Prodigy-HF Integration Release Notes
Explosion has released Prodigy-HF, a plugin that enables direct fine-tuning of Hugging Face transformer models on annotated data and the ability to upload datasets directly to the Hugging Face Hub.
Deploying Llama 2 on AWS Inferentia2 with optimum-neuron
Hugging Face has integrated optimum-neuron with the AWS Neuron SDK to enable the deployment of Llama 2 models on AWS Inferentia2 accelerators for high-performance text generation.