On this page
If you look at what developers are starring on GitHub this month, very little of it is a new coding agent. Claude Code and Codex lead the benchmarks, and the community is building around them: skills that carry know-how between tools, memory that survives a closed terminal, teams of agents that split a job, and reviewers that check what the agents wrote. I went through the projects behind the trend lists, checked their repositories, and kept what holds up. Star counts are from the GitHub API on October 10, 2026.
Where the agents stand#
The official Terminal-Bench 4.0 leaderboard, last updated on October 6, tests agents on real terminal tasks over 330 trials each. The top four:
| Agent | Model | Score |
|---|---|---|
| Claude Code | Opus 5.5 (max effort) | 64.8% ± 3.1 |
| Claude Code | Sonnet 5.5 (max effort) | 61.8% ± 2.9 |
| Codex | GPT-6 Astra (max effort) | 58.2% ± 2.8 |
| Codex | GPT-6.1 Sol | 58.2% ± 3.1 |
The margins of error overlap, so read this as two strong agents close together rather than a clear winner. Several trend posts quote these numbers from a vendor page; the leaderboard itself is the source to cite.
Skills became a portable standard#
I wrote about Claude Code skills last week: a folder with a SKILL.md that the agent loads only when a task matches. The news is that this format is now an open standard. The Agent Skills specification says it was originally developed by Anthropic and released for everyone, and its client list counts 46 tools, including Claude Code, Codex, Gemini CLI, GitHub Copilot, VS Code, Cursor, OpenHands and Goose.
So a skill you write once now works in most of the agents your team uses. The minimum is still two fields and some instructions:
---
name: release-notes
description: Write release notes from the pull requests merged since the last tag.
Use it when the user asks for release notes, a changelog entry or "what shipped".
---
# Release notes
1. Find the last tag with `git describe --tags --abbrev=0`.
2. List merged pull requests since then and group them by label.
3. Write one line per change, user-facing wording, no ticket numbers.Skill packs and harnesses#
The most-starred projects bundle many skills, hooks and settings into one install. addyosmani/agent-skills (about 104,000 stars) packages engineering skills for planning, building, testing, reviewing and shipping. ECC (about 276,000 stars) calls itself an agent harness: agents, skills, hooks, memory and a security scanner, tuned for Claude Code and synced to Codex.
npx skills add addyosmani/agent-skills
npx ecc-universal@2.2.3 setupCompressing what the agent reads#
Long sessions fail because the context fills with logs, test output and file dumps. Headroom (about 75,000 stars) compresses tool outputs, logs, retrieved chunks and files locally before they reach the model. Its README claims about 20 percent fewer tokens for coding agents, a figure the project measured itself.
pip install "headroom-ai[all]"
headroom wrap claudeMemory that survives the session#
Every session still starts from zero unless something writes down what happened. That is the gap memory tools fill, and it is the busiest category this month.
claude-mem (about 99,000 stars, version 13.35 released on October 9) records what the agent does during a session, compresses it with a model, and injects the relevant parts into later sessions. Despite the name it works with Codex, Gemini, Copilot and OpenCode too.
npx claude-mem installInside Claude Code you can install it as a plugin instead, with /plugin marketplace add thedotmack/claude-mem and then /plugin install claude-mem. One thing to know before you run the installer: it asks you to sign in to a hosted service with a 30-day trial. The README documents an opt-out, CLAUDE_MEM_ONLINE_OPTIN=false, if you want everything to stay on your machine.
mem0 (about 67,000 stars) is older, from 2023, and works at a different level: it is a memory layer you build into your own agents, keeping user, session and agent state. Install it with pip install mem0ai or npm install mem0ai. The benchmark scores in its README come from the hosted platform, and the README says the open-source version will not match them exactly.
Teams of agents#
OpenRig (about 6,600 stars) lets you define a team of agents in YAML and start it with one command, with Claude Code and Codex in the same team. It ships ready-made teams of two to seven agents, needs Node 22 or 24 and tmux, and runs on macOS and Linux, or Windows through WSL2.
npm install -g @openrig/cli
rig setup --dry-runStart with the dry run. OpenRig writes hooks and trust settings into your agents' configuration, and one optional bundle turns off permission prompts. Leave that bundle out until you know what each agent in the team is allowed to do.
On the model side, trend lists this month repeat that Kimi K2.6 runs "agent swarms" of 1,000 agents. The official model card says 300 sub-agents and 4,000 coordinated steps, and the model came out in April, not this month.
AI code review that mixes rules with a model#
open-code-review from Alibaba (about 46,000 stars, created in May) started as Alibaba's internal reviewer. Its design is the interesting part: ordinary code picks the files, bundles the context, matches rules and places the comments, and the model only does the analysis. That keeps comments on the right lines and costs fewer tokens than letting an agent explore the repository.
npm install -g @alibaba-group/open-code-review
ocr config provider
ocr reviewIt reads your git diff, posts comments line by line, and has plugins for Claude Code, Codex, Cursor and OpenCode. Its README claims higher precision than Claude Code at about a ninth of the tokens, measured by the project itself, and admits its recall is lower: it finds fewer issues, but more of them are real.
Sandboxes for agents#
The last piece is containment. Microsoft made Microsoft Execution Containers generally available on Windows 11 on October 7: a policy says which files and networks an agent may reach, and the operating system enforces it. Copilot, Codex, Replit and LM Studio support it today, and Claude Code support is announced. The weekly roundup has the details.
What I would try first#
- Write your own skills before installing packs. A ten-line skill for a task you repeat beats a hundred generic ones, and it now works across tools.
- Add memory on a side project. Watch what it stores for a week before you trust it with a work repository.
- Use a rule-based reviewer as a second opinion, not a gate. Low recall means a clean review does not mean clean code.
- Skip agent teams until one agent is reliable for you. Three agents multiply mistakes as easily as output.
Frequently asked questions#
What is the Agent Skills standard?
An open specification at agentskills.io for packaging agent instructions as a folder with a SKILL.md file, plus optional scripts, references and assets. It started at Anthropic and is supported by 46 clients, including Claude Code, Codex, Gemini CLI, GitHub Copilot and Cursor.
Which AI coding agent scores best on Terminal-Bench 4.0?
As of October 6, 2026, Claude Code with Opus 5.5 leads at 64.8 percent, followed by Claude Code with Sonnet 5.5 at 61.8 percent and Codex with GPT-6 models at 58.2 percent. The confidence intervals of the top two overlap.
Does claude-mem only work with Claude Code?
No. It works with Codex, Gemini, Copilot, OpenCode and other agents. Its installer offers a hosted service with a trial; the README documents how to opt out and keep data local.
Is AI code review good enough to replace human review?
Not yet. Tools like open-code-review aim for precise comments and admit lower recall, so they miss issues. Use them to catch the obvious problems before a person reviews the change.
Comments
No comments yet. Be the first.