✷ The archive · 11 labs · 450 dispatches
The labs
No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.
Collective Constitutional AI: Aligning a Language Model with Public Input
Anthropic and the Collective Intelligence Project developed a method to align a language model using a constitution drafted by approximately 1,000 members of the American public, resulting in a model with lower social bias than one aligned with an internal corporate constitution.
Anthropic Research: Decomposing Language Models With Dictionary Learning
Anthropic researchers have developed a method using dictionary learning to decompose language model layers into monosemantic features, allowing for the interpretation of complex neural network activations that are otherwise invisible at the individual neuron level.
Anthropic Decomposing Language Models Into Understandable Components
Anthropic researchers have developed a method using dictionary learning to decompose neural network activations into interpretable features, moving beyond the limitation of uninterpretable individual neurons.
Anthropic: Challenges in Evaluating AI Systems
Anthropic outlines the technical and operational difficulties in building robust AI evaluations, arguing that effective AI governance depends on overcoming these measurement challenges.
Anthropic Announces $4 Billion Amazon Investment and Expanded Claude 2 Access via AWS Bedrock
Anthropic announced a up‑to‑$4 billion investment from Amazon, making AWS the primary cloud for its models and expanding Claude 2 availability on Amazon Bedrock to bring safer, high‑performing AI to enterprises.
Anthropic Prompt Engineering Guide for Claude's 100k Token Context Window
Anthropic released a quantitative study showing that extracting reference quotes and providing multiple in-context examples dramatically improve Claude's recall on 70‑95k token documents.
Anthropic Responsible Scaling Policy
Anthropic has introduced a Responsible Scaling Policy (RSP) that uses AI Safety Levels (ASL) to mandate stricter safety and security protocols as AI models increase in capability and catastrophic risk.
Anthropic and BCG Partnership for Enterprise AI Deployment
Anthropic has partnered with Boston Consulting Group (BCG) to integrate Claude AI into enterprise strategic offerings, focusing on the deployment of safe and responsible generative AI solutions.
Claude Pro Release Notes
Anthropic has introduced Claude Pro, a paid subscription plan for Claude.ai that provides 5x more usage of the Claude 2 model, priority access, and early feature access for $20 (US) or 18 (UK) per month.
Anthropic and SK Telecom Partnership Announcement
Anthropic and SK Telecom (SKT) have entered a commercial partnership and strategic investment to develop a customized LLM for the telecommunications industry.
Claude Instant 1.2 Release Notes
Anthropic has released Claude Instant 1.2, a faster and lower-priced model that improves upon version 1.1 in math, coding, reasoning, and safety.
Studying Large Language Model Generalization with Influence Functions
Anthropic researchers use the EK-FAC approximation to scale influence functions to LLMs with up to 52 billion parameters, enabling the identification of specific training examples that drive model behavior.
Anthropic Research: Tracing Model Outputs to Training Data using Influence Functions
Anthropic researchers scaled influence functions to LLMs up to 52 billion parameters, discovering that model generalization becomes more abstract and conceptually driven as model scale increases.
Anthropic Frontier Threats Red Teaming for AI Safety
Anthropic has established a frontier threats red teaming framework to identify and mitigate national security risks, specifically finding that unmitigated LLMs could accelerate biological weaponization efforts within the next two to three years.
Anthropic Frontier Model Security Framework
Anthropic proposes a security framework for frontier AI models based on multi-party authorization and secure software development standards to prevent theft and misuse.
Anthropic Shows Question Decomposition Boosts Faithfulness of Model-Generated Reasoning
Anthropic announced that decomposing questions into subquestions markedly improves the faithfulness of large language model reasoning compared to standard chain‑of‑thought prompting.
Anthropic Study on Faithfulness of Chain-of-Thought Reasoning
Anthropic found that larger language models often generate unfaithful chain-of-thought explanations, with faithfulness varying widely across tasks and model sizes.
Claude 2 Release Notes
Anthropic has released Claude 2, featuring a 100K token context window, improved performance in coding and reasoning, and enhanced safety compared to Claude 1.3.
Anthropic Introduces GlobalOpinionQA Framework to Measure Subjective Global Opinion Representation in LLMs
Anthropic released GlobalOpinionQA, a dataset and quantitative framework that reveals large language models often mirror US and certain European opinions, and shows how prompting or translation can shift but not fully correct these biases.
Anthropic AI Accountability Recommendations for NTIA
Anthropic has proposed a comprehensive framework for AI accountability to the NTIA, focusing on standardized evaluations, risk-responsive assessments, and pre-registration of large training runs.
Anthropic Interpretability Dreams Research Vision
Anthropic outlines its long-term vision for mechanistic interpretability, focusing on resolving the challenge of superposition to enable the analysis of massive neural networks.
Anthropic Circuits Updates May 2023
Anthropic's May 2023 interpretability updates cover research into multiagent system failures, worker retraining program evidence, and a research version of Claude's progress on the Riemann zeta function.
Anthropic raises $450M Series C to scale reliable AI products
Anthropic announced a $450 million Series C round led by Spark Capital to accelerate development of its Claude assistant, expand product offerings, and fund AI safety research.
Anthropic and Zoom Partnership and Investment
Anthropic and Zoom have partnered to integrate Claude AI into Zoom's enterprise collaboration products and Zoom Ventures has invested in Anthropic.
Anthropic Announces 100K Token Context Windows for Claude
Anthropic expanded Claude's context window from 9K to 100K tokens, enabling the model to ingest and analyze hundreds of pages of text in under a minute.
Claude's Constitution: Implementing Constitutional AI for Model Alignment
Anthropic introduces Constitutional AI, a method for training Claude to be helpful, honest, and harmless using an explicit set of written principles rather than relying solely on implicit human feedback.
Anthropic Research: Distributed Representations, Composition, and Superposition
Anthropic clarifies the distinction between composition and superposition in distributed representations, explaining how these two mechanisms impact generalization and linear computability in neural networks.
Anthropic and Scale Partnership for Enterprise Generative AI
Anthropic has partnered with Scale to integrate the Claude AI assistant into Scale's platform, providing enterprises with deployment tools, prompt engineering, and secure data integration.
Anthropic Proposal for Increased NIST Funding for AI Measurement
Anthropic proposes ambitiously funding the National Institute of Standards and Technology (NIST) to develop standardized AI measurement tools and safety thresholds, which they argue is a prerequisite for effective AI regulation.
Anthropic Reveals Privileged Bases in Transformer Residual Streams
Anthropic discovered that transformer residual stream dimensions are not arbitrary but align with privileged bases, likely due to Adam's per-dimension normalizers, challenging prior theoretical assumptions.
Introducing Claude
Anthropic has released Claude, a next-generation AI assistant designed to be helpful, honest, and harmless, available in high-performance and fast, lightweight versions.
Anthropic's Core Views on AI Safety
Anthropic outlines an empirically-driven, portfolio-based approach to AI safety to mitigate catastrophic risks associated with the rapid scaling of transformative AI systems.
Anthropic Research: The Capacity for Moral Self-Correction in LLMs
Anthropic research demonstrates that Large Language Models (LLMs) trained with RLHF can morally self-correct to avoid harmful outputs when instructed, a capability that emerges at 22B parameters.
Anthropic Partners with Google Cloud for AI Infrastructure
Anthropic has selected Google Cloud as its cloud provider to leverage GPU and TPU clusters for training, scaling, and deploying its AI systems, including the Claude assistant.
Anthropic Research: Superposition, Memorization, and Double Descent
Anthropic researchers investigate the relationship between superposition, memorization, and overfitting in simple neural networks, suggesting that superposition allows models to store more features than they have neurons during memorization.
Anthropic Research: Discovering Language Model Behaviors with Model-Written Evaluations
Anthropic researchers developed a method to automatically generate high-quality evaluations using language models to discover novel behaviors, including inverse scaling and sycophancy in larger models.
Constitutional AI: Harmlessness from AI Feedback
Anthropic introduces Constitutional AI, a method for training harmless AI assistants using a set of rules (a constitution) and AI feedback instead of human labels to reduce harmful outputs.
Measuring Progress on Scalable Oversight for Large Language Models
Anthropic researchers propose a framework for empirically studying scalable oversight to ensure humans can supervise AI systems that may eventually outperform them on complex tasks.
Anthropic Toy Models of Superposition Research
Anthropic researchers use small ReLU networks to demonstrate that models can represent more features than they have dimensions through a phenomenon called superposition, provided those features are sparse.
Anthropic Red Teaming Language Models Report – Methods, Scaling Behaviors, and Lessons Learned
Anthropic released a detailed study showing that reinforcement‑learning‑from‑human‑feedback (RLHF) models become harder to red‑team as they scale, while other model types show flat red‑teamability, and they published a 38,961‑attack dataset to help the community improve safety.
Anthropic Research: Language Models (Mostly) Know What They Know
Anthropic research demonstrates that large language models can effectively evaluate the validity of their own claims and predict their likelihood of knowing an answer, providing a path toward more honest AI.
Anthropic Softmax Linear Units (SoLU) Research
Anthropic introduces Softmax Linear Units (SoLU), an architectural change to MLP activation functions that increases the fraction of interpretable neurons without sacrificing model performance.
Anthropic Research: Scaling Laws and Interpretability of Learning from Repeated Data
Anthropic researchers found that repeating a small fraction of training data can severely degrade LLM performance by consuming model capacity for memorization and damaging generalization structures like induction heads.
Anthropic Series B Funding for AI Safety and Research
Anthropic has raised $580 million in Series B funding to build large-scale experimental infrastructure aimed at improving the steerability, interpretability, and robustness of computationally intensive AI models.
Anthropic: Training a Helpful and Harmless Assistant with RLHF
Anthropic demonstrates that Reinforcement Learning from Human Feedback (RLHF) improves language model performance across most NLP evaluations while ensuring the assistant remains helpful and harmless.
Anthropic In-context Learning and Induction Heads
Anthropic announced a research post titled “In-context Learning and Induction Heads” on March 8 2022, but the page contains only related links and no substantive technical details.
Anthropic Research: Predictability and Surprise in Large Generative Models
Anthropic identifies a tension between the predictable scaling of loss in large generative models and the unpredictable emergence of specific capabilities and outputs, creating challenges for AI safety and policy.
Anthropic Announces Mathematical Framework for Transformer Circuits
Anthropic released a brief announcement of a new mathematical framework for transformer circuits, highlighting its potential to deepen understanding of model behavior but providing no technical details in the post.
Anthropic Research: A General Language Assistant as a Laboratory for Alignment
Anthropic explores methods to create a helpful, honest, and harmless general-purpose language assistant, finding that ranked preference modeling scales more effectively than imitation learning or binary discrimination.
Anthropic Series A Funding for Reliable General AI Systems
Anthropic has raised $124 million in Series A funding to develop large-scale AI systems that are steerable, interpretable, and robust.