Many of the cheap, fast techniques now used to steer large language models get the job done at a steep cost to the fluency of what the models write, Apple researchers reported on Thursday.

In a study published on Apple’s machine learning research site, the authors examined a range of conditioning methods in two settings: injection, in which a model is made to express a target concept, and removal, in which a concept is taken out. Such methods, they wrote, are usually judged on a narrow question, whether the concept takes hold, and generation quality gets little attention.

Efficient steering methods, which adjust a model’s internal activations as it generates text, paid the highest price. “We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency,” the researchers wrote.

The study also identified an interaction the authors said had been overlooked: activation steering worked far less well on instruction-tuned models than on their base counterparts. The finding bears on current practice, because the models behind most widely used chatbots are instruction-tuned.

Simple prompting and full supervised fine-tuning, by contrast, remained viable options for injecting a concept, the authors found. Neither was as good at removing one.

The study offers a cheaper way to run such evaluations. Text-based metrics that cost little to compute correlated highly with scores from LLM-as-judge grading, in which one model rates another’s output, and they shed light on how each conditioning method behaves, according to the paper.

Its authors are Iuri Macocco and Marco Baroni of Universitat Pompeu Fabra, working with Pau Rodríguez Lopez, Arno Blaas, Luca Zappella and Xavier Suau Cuadros.