Google employees have been trying out an internal version of the company’s Gemini 4 model, code-named Carbon, and at least one of them said it codes about as well as Anthropic’s top model, Business Insider reported on Friday.

Carbon showed up in recent days on Jetski, Google’s internal coding platform, according to documents and screenshots seen by the Business Insider reporter Hugh Langley. It may be a checkpoint, meaning an update, to Argon, the Gemini 4 model Google announced on Sept. 30. Most of the public still can’t use Argon.

One employee told Business Insider that Carbon “feels like Opus 5.5” for coding but said it would need more testing. Opus 5.5 is Anthropic’s most powerful model for long-running agentic coding work.

The internal message threads read the same way. “Feels comparable to Opus 5.5,” one employee wrote about the latest update. “Carbon is really good!” wrote another. A third posted, “New Gemini pro next model feels kinda nice,” a reference to the “Next” label Google usually attaches to Gemini updates inside the company.

Staff were mostly positive about the early Argon builds too, according to employees and internal threads. Not everyone. One employee said those versions fell short on some coding tasks and compared them to Claude Opus 5, an older Anthropic release.

Google has tested a run of Gemini 4 models under the internal names Argon, Barium and Carbon, according to an internal document. The code names don’t line up with what the public sees. One document said a model called “Barium-B” was picked to become the public Argon. It was not clear whether Carbon would ship as another Argon update or as a separate model in the Gemini 4 family, and nothing guarantees Google will release it at all. Google declined to comment.

Coding is where Google has the most to prove. Anthropic and OpenAI have pulled ahead with developers on models built for complex engineering work. In its Sept. 30 announcement, Google said Argon set a new state of the art on the DeepSWE v1.1 software benchmark with a score of 77.9 percent and tied for first on CWE-bench v1, a test of fixing security flaws, at 68 percent. Google also cited strong results on legal and finance work. Even so, Argon trailed in five coding and terminal rows of Google’s own comparison table, where Claude Opus 5.5 and OpenAI’s GPT-6 Astra led, according to a tally by the tracker LLM Stats.

Argon’s release has been narrow. Google built the model with a heavy focus on defensive cybersecurity and is giving it first to partners in its Fairwind Program, an initiative for testing cyber vulnerabilities in new models. Those defenders and Google’s internal teams will also get a version without cyber guardrails, SecurityWeek reported.

Google said it would open Argon to developers, businesses and consumers, “starting with paid API customers and Google AI Ultra subscribers,” and described the plan as “a phased approach.” It set an introductory price of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after that, TechCrunch reported along with the company’s post.

The internal tests come during a busy stretch for the company. This week it introduced a “universal” Gemini agent for the workplace as the big AI labs compete over agentic assistants. Ten days after the Argon announcement, Google’s post still gave no date for the wider release. It said only: “Rolling out soon.”