---
title: "Apple Study Finds Cheap Methods for Steering A.I. Models Exact a Steep Cost in Fluency"
description: "In a systematic study, activation steering also worked far worse on instruction-tuned models than on base models."
author: "rews desk"
published: 2026-10-01T08:50:25.355Z
modified: 2026-10-01T22:11:20Z
url: https://rews.cc/a/apple-study-finds-cheap-methods-for-steering-a-i-models-exac-03c9af
language: en
tags: ["ai", "llm", "openai", "regulation", "automation", "tech"]
publisher: "Rews (https://rews.cc)"
---

# Apple Study Finds Cheap Methods for Steering A.I. Models Exact a Steep Cost in Fluency

*In a systematic study, activation steering also worked far worse on instruction-tuned models than on base models.*

By rews desk · October 1, 2026 · https://rews.cc/a/apple-study-finds-cheap-methods-for-steering-a-i-models-exac-03c9af

## In brief

- Apple researchers published a systematic study of methods for conditioning large language models
- The study tested methods for injecting a concept into model output and for removing one
- Activation steering was far less effective on instruction-tuned models than on base models
- Prompting and supervised fine-tuning injected concepts well but removed them poorly
- Cheap text-based metrics correlated highly with costly LLM-as-judge scores

Many of the cheap, fast techniques now used to steer large language models get the job done at a steep cost to the fluency of what the models write, Apple researchers reported on Thursday.

In a study published on Apple’s machine learning research site, the authors examined a range of conditioning methods in two settings: injection, in which a model is made to express a target concept, and removal, in which a concept is taken out. Such methods, they wrote, are usually judged on a narrow question, whether the concept takes hold, and generation quality gets little attention.

Efficient steering methods, which adjust a model’s internal activations as it generates text, paid the highest price. “We find that efficient steering methods frequently achieve conditioning at a steep cost to fluency,” the researchers wrote.

The study also identified an interaction the authors said had been overlooked: activation steering worked far less well on instruction-tuned models than on their base counterparts. The finding bears on current practice, because the models behind most widely used chatbots are instruction-tuned.

Simple prompting and full supervised fine-tuning, by contrast, remained viable options for injecting a concept, the authors found. Neither was as good at removing one.

The study offers a cheaper way to run such evaluations. Text-based metrics that cost little to compute correlated highly with scores from LLM-as-judge grading, in which one model rates another’s output, and they shed light on how each conditioning method behaves, according to the paper.

Its authors are Iuri Macocco and Marco Baroni of Universitat Pompeu Fabra, working with Pau Rodríguez Lopez, Arno Blaas, Luca Zappella and Xavier Suau Cuadros.
