Contrastive Language Models (CLM) for Fast Agentic Decision-Making
Contrastive Language Models (CLMs) enable high-speed agentic decision-making by treating action selection as a retrieval problem rather than a generative one. By utilizing a contrastive learning objective to align state and action embeddings in a shared space, CLM-8B achieves performance comparable to Jev across computer-use, gaming, and tool-calling tasks while reducing latency by up to 9x.
Architecture: Disaggregated State and Action Encoders
CLM replaces the standard autoregressive (AR) generation process with a dual-encoder architecture. The system consists of a state encoder and an action encoder, both utilizing a frozen LLM backbone followed by a trainable MLP projection head (approximately 20M parameters).
- Mechanism: Both encoders map their respective inputs into a shared embedding space. The score for a state-action pair is calculated as the cosine similarity between their embeddings.
- Inference Efficiency: Because states and actions are disaggregated, their embeddings can be cached independently. In environments where the action set is fixed (e.g., Super Mario), the model only needs to recompute the state embedding at each step, reusing cached action embeddings. This reduces the computational cost from multiple forward passes to a single pass per step.
- Scaling: At approximately 1,000 candidate actions, CLM is reported to be 13x faster than Jev.
Training Pipeline and Scaling Laws
CLM is trained in three distinct stages to progressively refine state-action alignment:
- Pre-training: The model is trained on ~60M Nemotron Q&A pairs using a bidirectional InfoNCE loss, where questions are treated as states and answers as actions. This establishes broad semantic representations.
- Mid-training: The model is exposed to ~30M synthetic hard negatives generated by Gemini 2.5 Flash-Lite. These are semantically similar but incorrect answers designed to improve fine-grained discrimination.
- Post-training: The model is trained on ~1M agent trajectories from the Agent Data Protocol (ADP) dataset and terminal traces from Endless-Terminals and LiteCoder-Terminal-SFT to adapt the representation space for agentic environments.
To prevent catastrophic forgetting during post-training, the team uses co-training with data replay, mixing 40% Nemotron DQA examples with 60% agentic trajectories.
Scaling Laws: The test contrastive loss follows a power law relative to training compute, dataset size, projection-head size, and encoder size. Scaling the encoder size provides the strongest performance gains. The optimal projection head size ($N^*$) grows linearly with the number of training tokens ($D^{1.02}$), averaging roughly 310 tokens per parameter.
Performance and Benchmarks
CLM-8B demonstrates strong verification capabilities, particularly when used as a reward model to select the best solution from multiple candidates sampled by other models (e.g., Opus 5 or Fable 5).
- Agentic Coding: CLM achieves state-of-the-art (SOTA) results on DeepSWE (81.6%) and Terminal Bench 2.1 (87.6%).
- Latency: On H100 GPUs, CLM delivers 4-6x faster inference than Jev for verification tasks.
- Generalization: Pre-training reshapes the model's representations into a useful decision space, allowing it to correctly rank gold answers higher than baseline LLM embeddings (e.g., Qwen3-8B).
Community Insights and Technical Critique
Discussion among technical users highlights several critical perspectives on the CLM approach:
- Verification vs. Generation: Some users note that the high scores on DeepSWE are achieved by acting as a verifier (selecting the best of several candidates) rather than generating the solution from scratch.
- Inductive Bias: Critics argue that while cosine similarity over independently encoded vectors is effective for large-K semantic action matching (like WikiRacing), it may struggle with "typed-decision" loads such as date arithmetic or negation chains.
- Terminology: There is pushback against the use of the term "System One" to describe the model, with users arguing that the Kahneman framework describes human cognitive processes rather than a speed-based classification for AI models.
- Hardware Efficiency: The ability to train projection heads in about an hour on a single RTX 4090 makes this architecture particularly attractive for developers working with consumer-grade hardware.
"I think this approach is really cool, and it does suggest that one might be able to use a modern LLM to process an input and then extract the model’s next agentic step in a very fast, non-AR manner, with results comparably good to the usual AR decoding."
Future Roadmap
The project is expanding toward a multimodal CLM-35B, which is currently in training. Future goals include extending the architecture to support images and video for robotics and computer-use tasks, as well as further expanding the data recipe for hard-negative mining and agentic post-training.
Sources
Related
- Dispatch
- Project
- Project
- Project
- Dispatch