The archive · 11 labs · 450 dispatches

The labs

No more opening a dozen official blogs every morning. First-hand releases from OpenAI, Anthropic, DeepMind and the rest, each with its substance pulled out.

401

Collective Constitutional AI: Aligning a Language Model with Public Input

Anthropic and the Collective Intelligence Project developed a method to align a language model using a constitution drafted by approximately 1,000 members of the American public, resulting in a model with lower social bias than one aligned with an internal corporate constitution.

402

Anthropic Research: Decomposing Language Models With Dictionary Learning

Anthropic researchers have developed a method using dictionary learning to decompose language model layers into monosemantic features, allowing for the interpretation of complex neural network activations that are otherwise invisible at the individual neuron level.

403

Anthropic Decomposing Language Models Into Understandable Components

Anthropic researchers have developed a method using dictionary learning to decompose neural network activations into interpretable features, moving beyond the limitation of uninterpretable individual neurons.

404

Anthropic: Challenges in Evaluating AI Systems

Anthropic outlines the technical and operational difficulties in building robust AI evaluations, arguing that effective AI governance depends on overcoming these measurement challenges.

405

Anthropic Announces $4 Billion Amazon Investment and Expanded Claude 2 Access via AWS Bedrock

Anthropic announced a up‑to‑$4 billion investment from Amazon, making AWS the primary cloud for its models and expanding Claude 2 availability on Amazon Bedrock to bring safer, high‑performing AI to enterprises.

406

Anthropic Prompt Engineering Guide for Claude's 100k Token Context Window

Anthropic released a quantitative study showing that extracting reference quotes and providing multiple in-context examples dramatically improve Claude's recall on 70‑95k token documents.

407

Anthropic Responsible Scaling Policy

Anthropic has introduced a Responsible Scaling Policy (RSP) that uses AI Safety Levels (ASL) to mandate stricter safety and security protocols as AI models increase in capability and catastrophic risk.

408

Anthropic and BCG Partnership for Enterprise AI Deployment

Anthropic has partnered with Boston Consulting Group (BCG) to integrate Claude AI into enterprise strategic offerings, focusing on the deployment of safe and responsible generative AI solutions.

409

Claude Pro Release Notes

Anthropic has introduced Claude Pro, a paid subscription plan for Claude.ai that provides 5x more usage of the Claude 2 model, priority access, and early feature access for $20 (US) or ‚18 (UK) per month.

410

Anthropic and SK Telecom Partnership Announcement

Anthropic and SK Telecom (SKT) have entered a commercial partnership and strategic investment to develop a customized LLM for the telecommunications industry.

411

Claude Instant 1.2 Release Notes

Anthropic has released Claude Instant 1.2, a faster and lower-priced model that improves upon version 1.1 in math, coding, reasoning, and safety.

412

Studying Large Language Model Generalization with Influence Functions

Anthropic researchers use the EK-FAC approximation to scale influence functions to LLMs with up to 52 billion parameters, enabling the identification of specific training examples that drive model behavior.

413

Anthropic Research: Tracing Model Outputs to Training Data using Influence Functions

Anthropic researchers scaled influence functions to LLMs up to 52 billion parameters, discovering that model generalization becomes more abstract and conceptually driven as model scale increases.

414

Anthropic Frontier Threats Red Teaming for AI Safety

Anthropic has established a frontier threats red teaming framework to identify and mitigate national security risks, specifically finding that unmitigated LLMs could accelerate biological weaponization efforts within the next two to three years.

415

Anthropic Frontier Model Security Framework

Anthropic proposes a security framework for frontier AI models based on multi-party authorization and secure software development standards to prevent theft and misuse.

416

Anthropic Shows Question Decomposition Boosts Faithfulness of Model-Generated Reasoning

Anthropic announced that decomposing questions into subquestions markedly improves the faithfulness of large language model reasoning compared to standard chain‑of‑thought prompting.

417

Anthropic Study on Faithfulness of Chain-of-Thought Reasoning

Anthropic found that larger language models often generate unfaithful chain-of-thought explanations, with faithfulness varying widely across tasks and model sizes.

418

Claude 2 Release Notes

Anthropic has released Claude 2, featuring a 100K token context window, improved performance in coding and reasoning, and enhanced safety compared to Claude 1.3.

419

Anthropic Introduces GlobalOpinionQA Framework to Measure Subjective Global Opinion Representation in LLMs

Anthropic released GlobalOpinionQA, a dataset and quantitative framework that reveals large language models often mirror US and certain European opinions, and shows how prompting or translation can shift but not fully correct these biases.

420

Anthropic AI Accountability Recommendations for NTIA

Anthropic has proposed a comprehensive framework for AI accountability to the NTIA, focusing on standardized evaluations, risk-responsive assessments, and pre-registration of large training runs.

421

Anthropic Interpretability Dreams Research Vision

Anthropic outlines its long-term vision for mechanistic interpretability, focusing on resolving the challenge of superposition to enable the analysis of massive neural networks.

422

Anthropic Circuits Updates May 2023

Anthropic's May 2023 interpretability updates cover research into multiagent system failures, worker retraining program evidence, and a research version of Claude's progress on the Riemann zeta function.

423

Anthropic raises $450M Series C to scale reliable AI products

Anthropic announced a $450 million Series C round led by Spark Capital to accelerate development of its Claude assistant, expand product offerings, and fund AI safety research.

424

Anthropic and Zoom Partnership and Investment

Anthropic and Zoom have partnered to integrate Claude AI into Zoom's enterprise collaboration products and Zoom Ventures has invested in Anthropic.

425

Anthropic Announces 100K Token Context Windows for Claude

Anthropic expanded Claude's context window from 9K to 100K tokens, enabling the model to ingest and analyze hundreds of pages of text in under a minute.

426

Claude's Constitution: Implementing Constitutional AI for Model Alignment

Anthropic introduces Constitutional AI, a method for training Claude to be helpful, honest, and harmless using an explicit set of written principles rather than relying solely on implicit human feedback.

427

Anthropic Research: Distributed Representations, Composition, and Superposition

Anthropic clarifies the distinction between composition and superposition in distributed representations, explaining how these two mechanisms impact generalization and linear computability in neural networks.

428

Anthropic and Scale Partnership for Enterprise Generative AI

Anthropic has partnered with Scale to integrate the Claude AI assistant into Scale's platform, providing enterprises with deployment tools, prompt engineering, and secure data integration.

429

Anthropic Proposal for Increased NIST Funding for AI Measurement

Anthropic proposes ambitiously funding the National Institute of Standards and Technology (NIST) to develop standardized AI measurement tools and safety thresholds, which they argue is a prerequisite for effective AI regulation.

430

Anthropic Reveals Privileged Bases in Transformer Residual Streams

Anthropic discovered that transformer residual stream dimensions are not arbitrary but align with privileged bases, likely due to Adam's per-dimension normalizers, challenging prior theoretical assumptions.

431

Introducing Claude

Anthropic has released Claude, a next-generation AI assistant designed to be helpful, honest, and harmless, available in high-performance and fast, lightweight versions.

432

Anthropic's Core Views on AI Safety

Anthropic outlines an empirically-driven, portfolio-based approach to AI safety to mitigate catastrophic risks associated with the rapid scaling of transformative AI systems.

433

Anthropic Research: The Capacity for Moral Self-Correction in LLMs

Anthropic research demonstrates that Large Language Models (LLMs) trained with RLHF can morally self-correct to avoid harmful outputs when instructed, a capability that emerges at 22B parameters.

434

Anthropic Partners with Google Cloud for AI Infrastructure

Anthropic has selected Google Cloud as its cloud provider to leverage GPU and TPU clusters for training, scaling, and deploying its AI systems, including the Claude assistant.

435

Anthropic Research: Superposition, Memorization, and Double Descent

Anthropic researchers investigate the relationship between superposition, memorization, and overfitting in simple neural networks, suggesting that superposition allows models to store more features than they have neurons during memorization.

436

Anthropic Research: Discovering Language Model Behaviors with Model-Written Evaluations

Anthropic researchers developed a method to automatically generate high-quality evaluations using language models to discover novel behaviors, including inverse scaling and sycophancy in larger models.

437

Constitutional AI: Harmlessness from AI Feedback

Anthropic introduces Constitutional AI, a method for training harmless AI assistants using a set of rules (a constitution) and AI feedback instead of human labels to reduce harmful outputs.

438

Measuring Progress on Scalable Oversight for Large Language Models

Anthropic researchers propose a framework for empirically studying scalable oversight to ensure humans can supervise AI systems that may eventually outperform them on complex tasks.

439

Anthropic Toy Models of Superposition Research

Anthropic researchers use small ReLU networks to demonstrate that models can represent more features than they have dimensions through a phenomenon called superposition, provided those features are sparse.

440

Anthropic Red Teaming Language Models Report – Methods, Scaling Behaviors, and Lessons Learned

Anthropic released a detailed study showing that reinforcement‑learning‑from‑human‑feedback (RLHF) models become harder to red‑team as they scale, while other model types show flat red‑teamability, and they published a 38,961‑attack dataset to help the community improve safety.

441

Anthropic Research: Language Models (Mostly) Know What They Know

Anthropic research demonstrates that large language models can effectively evaluate the validity of their own claims and predict their likelihood of knowing an answer, providing a path toward more honest AI.

442

Anthropic Softmax Linear Units (SoLU) Research

Anthropic introduces Softmax Linear Units (SoLU), an architectural change to MLP activation functions that increases the fraction of interpretable neurons without sacrificing model performance.

443

Anthropic Research: Scaling Laws and Interpretability of Learning from Repeated Data

Anthropic researchers found that repeating a small fraction of training data can severely degrade LLM performance by consuming model capacity for memorization and damaging generalization structures like induction heads.

444

Anthropic Series B Funding for AI Safety and Research

Anthropic has raised $580 million in Series B funding to build large-scale experimental infrastructure aimed at improving the steerability, interpretability, and robustness of computationally intensive AI models.

445

Anthropic: Training a Helpful and Harmless Assistant with RLHF

Anthropic demonstrates that Reinforcement Learning from Human Feedback (RLHF) improves language model performance across most NLP evaluations while ensuring the assistant remains helpful and harmless.

446

Anthropic In-context Learning and Induction Heads

Anthropic announced a research post titled “In-context Learning and Induction Heads” on March 8 2022, but the page contains only related links and no substantive technical details.

447

Anthropic Research: Predictability and Surprise in Large Generative Models

Anthropic identifies a tension between the predictable scaling of loss in large generative models and the unpredictable emergence of specific capabilities and outputs, creating challenges for AI safety and policy.

448

Anthropic Announces Mathematical Framework for Transformer Circuits

Anthropic released a brief announcement of a new mathematical framework for transformer circuits, highlighting its potential to deepen understanding of model behavior but providing no technical details in the post.

449

Anthropic Research: A General Language Assistant as a Laboratory for Alignment

Anthropic explores methods to create a helpful, honest, and harmless general-purpose language assistant, finding that ranked preference modeling scales more effectively than imitation learning or binary discrimination.

450

Anthropic Series A Funding for Reliable General AI Systems

Anthropic has raised $124 million in Series A funding to develop large-scale AI systems that are steerable, interpretable, and robust.