

A substantial improvement in intelligence and behavior over Composer 2, particularly on long-horizon agentic tasks.
Loading comments…
Achievement
Project Info
Product Keywords
Composer 2.5 is a major update to the AI coding assistant available in Cursor. It represents a substantial leap in intelligence and behavior over its predecessor, Composer 2, with particular strength in long-horizon agentic tasks. The model is built on the same open-source checkpoint as Composer 2—Moonshot's Kimi K2.5—but benefits from scaled training, more complex reinforcement learning environments, and new learning methods that improve both capability and usability.
Composer 2.5 introduces a novel training technique that provides localized feedback at specific points in a rollout. Instead of relying solely on a final reward signal—which can be noisy over hundreds of thousands of tokens—the model receives hints inserted directly into the context where a behavior needs improvement. This allows Composer 2.5 to learn from mistakes like bad tool calls or confusing explanations without being penalized for the entire trajectory.
The model is trained on 25 times more synthetic tasks than Composer 2. These tasks are dynamically created and selected throughout the training run, ensuring the model continues to face harder problems as its coding ability improves. Approaches like feature deletion—where the agent must remove code while keeping a test suite passing—ground the synthetic data in real-world codebase challenges.
Beyond raw coding benchmarks, Composer 2.5 has been refined on behavioral dimensions that matter for real-world use. The model communicates more clearly, calibrates its effort appropriately to the task, and is generally more pleasant to collaborate with over long sessions.
"Composer 2.5 is better at sustained work on long-running tasks, follows complex instructions more reliably, and is more pleasant to collaborate with."
This combination of sustained intelligence and behavioral polish is rare in AI coding assistants. While many models can handle short, well-defined tasks, Composer 2.5 excels at the kind of extended, multi-step work that defines real software development. The targeted feedback training method means it learns from specific mistakes rather than being punished for entire trajectories, making it both smarter and more adaptable in practice.
You're a developer who regularly works on complex, long-running coding tasks and wants an AI assistant that maintains focus, follows nuanced instructions, and communicates clearly throughout the process. If you've found other coding assistants lose coherence on multi-step work or fail to learn from localized mistakes, Composer 2.5 offers a meaningful upgrade in both intelligence and collaboration quality.
Other tools you might consider
Three things kept hitting me using Claude Code. Tab switching every 30 seconds just to check if it's still running. Claude silently blocking for 12 minutes while I was in another app. Coming back to find it finished or stuck 15 minutes ago. So I built CodeBreak. A pixel-art character walks your screen while CC runs. It celebrates when done, panics when it needs you, sulks on errors. $7 one-time. No sub. All future updates are free. Its for CC now, eventually will be Universal for all AI tools.
Whats 1Code? An app to run your Claude Code agents in parallel that works on Mac and Web. On Mac - run locally, with or without worktrees. On Web - run in remote sandboxes with live previews of your app, mobile included, so you can check on agents from anywhere. Running multiple Claude Codes in parallel dramatically sped up how we build features.
LobeHub is a Chief Agent Operator (CAO) that builds, runs, and coordinates your AI agent team. Describe a goal, and it assembles the right agents/skills, runs tasks in parallel in the cloud, routes work across models, and reports back only when decisions are needed—via your existing channels (Slack/Discord/Telegram/iMessage). Less tab-switching, more outcomes.
AI agents render UI slowly, expensively, inconsistently and inference bills balloon from it. Montage fixes it: emit a tiny intent schema, we compile production components server-side: 10x faster, 50-100x fewer tokens, model and framework agnostic. Now one M1 API call generates rich interactive visuals, hosts them as live UIs with persistent state, and styles to your brand. Don't let your agents reinvent UI every turn - ship them on Montage!
Maker
blueprint_b
Alternatives
Loading comments…