Researchers at Google published the first system for watermarking proteins designed by artificial intelligence on Wednesday, a tool they say could make it easier to spot biological threats before they are ever made.
The result, described in a paper in the journal Nature, embeds a hidden statistical mark in a protein’s amino acid sequence without changing what the protein does. Companies that synthesize DNA to order could use it to tell a novel protein from a trusted laboratory apart from one designed by nobody in particular.
That distinction has become a security problem. Software that designs proteins has produced enzymes that can digest plastics and molecules that block snake venom, but the same tools could, in principle, build toxins or alter the proteins of viruses. DNA synthesis companies screen orders against libraries of known threats, and a protein invented by A.I. often matches nothing in those libraries. When the risk was pointed out nearly a year ago, no one had a fix.
Proteins make a poor hiding place. Only 20 amino acids exist, and a change at a single position can shut a protein down. Five hundred amino acids counts as a large protein; a digital photograph offers millions of pixels in which to conceal a signal.
Google’s system, SynthIDBio, adapts SynthID, the watermarking technology the company built for A.I.-generated text and images. It works inside ProteinMPNN, one of the most widely used protein design programs, written in the laboratory of David Baker, who shared the Nobel Prize with the head of DeepMind, Google’s A.I. unit.
ProteinMPNN assembles a design one amino acid at a time along the protein’s backbone. At each step, SynthIDBio consults a secret key — similar to a cryptographic one — and the amino acids already chosen, then proposes the next one. The design software accepts the suggestion only if it still yields a working protein. Where a spot could equally well take leucine or valine, the watermark gets its way; where one particular amino acid is essential, it stays silent.
The mark ends up scattered along the full length of the protein, so detection is statistical rather than a yes-or-no test: scan the sequence with the key in hand and count how often the suggested amino acids turn up. Google wrote the software for that task too.
To check that the process did no damage, the team designed watermarked versions of proteins meant to bind natural targets that earlier A.I. designs had gone after. The marked proteins bound as intended. Binding is an easier test than engineering a catalyst, but the experiment produced no sign of trouble.
Google’s proposal is that trusted organizations, such as universities and large biotechnology companies, hand their keys to the DNA synthesizers. An unknown sequence that checks out as a trusted design could pass quickly, letting screeners spend their time on sequences that look A.I.-designed but belong to no one. The system tightens screening rather than guaranteeing it, the authors said.
The paper’s authors list the ways their own system could fail. Keys must be distributed and kept without leaking. Very short proteins may carry too few marked amino acids to detect. Padding a marked protein with an unmarked one — the paper raises a natural fluorescent protein as the example — could dilute the signal, and where the detection cutoff is set trades false alarms against misses.
And plenty of design tools are not ProteinMPNN. Some also place amino acids one at a time and could accept SynthIDBio the same way. Others work differently, and Google has not yet said how watermarking would reach them.

