Controlled Text Generation with Hidden Representation Transformations
Vaibhav Kumar, Hana Koorehdavoudi, Masud Moshtaghi, Amita Misra, Ankit Chadha, Emilio Ferrara
Controlled generationFindings of ACL 2023, pages 9440–9455 · Amazon Alexa AI; University of Southern California
PDF arXiv DOI Publisher page Code BibTeX
CHRT steers a language model away from toxic or negative text by editing its internal number vectors, cutting toxicity from 0.827 to 0.085 while adding only 0.01 seconds per generation.
Abstract
We propose CHRT (Control Hidden Representation Transformation) – a controlled language generation framework that steers large language models to generate text pertaining to certain attributes (such as toxicity). CHRT gains attribute control by modifying the hidden representation of the base model through learned transformations. We employ a contrastive-learning framework to learn these transformations that can be combined to gain multi-attribute control. The effectiveness of CHRT is experimentally shown by comparing it with seven baselines over three attributes. CHRT outperforms all the baselines in the task of detoxification, positive sentiment steering, and text simplification while minimizing the loss in linguistic qualities. Further, our approach has the lowest inference latency of only 0.01 seconds more than the base model, making it the most suitable for high-performance production environments. We open-source our code and release two novel datasets to further propel controlled language generation research.
Why it matters
Most ways to stop a language model from producing toxic text either slow generation down a lot or make the text read worse. CHRT trains a small block that rewrites the model's internal representations, so the base model and its decoder stay untouched at run time. On the paper's tests it cut toxicity further than five prior methods while adding 0.01 seconds per continuation, against 9.30 extra seconds for PPLM. Attribute blocks can also be mixed at inference time without retraining.
Key results
- On RealToxicityPrompts, average maximum toxicity falls from 0.827 for the base GPT-2 to 0.085 for CHRT12, below DExperts at 0.154, and human raters likewise score CHRT12 lowest at 0.027 against 0.369 for GPT-2.
- On the new RealSentimentPrompts set, average maximum negative sentiment drops from 0.934 for GPT-2 to 0.094 for CHRT12, which the paper states is more than 62% lower than DExperts at 0.249, with negative-sentiment probability falling from 0.534 to 0.005.
- On RealSimplicityPrompts, the probability of generating simple continuations rises from 0.259 for GPT-2 to 0.996 for CHRT12 with average maximum simplicity of 0.995, against 0.959 and 0.992 for DExperts.
- Generation time per 25-token continuation goes from 0.811 seconds for GPT-2 to 0.823 seconds for CHRT, an increase of +0.01, compared with 10.12 seconds for PPLM (+9.30) and 1.989 seconds for DExperts (+1.17).
- Null or negativeFluency gets worse, not better, on sentiment steering: CHRT12 perplexity rises from 17.372 for GPT-2 to 28.160, a 10.84 point loss that is also worse than NLC at 17.542 and DAPT at 19.570.
- Null or negativeIn the human evaluation, topicality shows no advantage for CHRT12 at 1.160 against 1.320 for GPT-2, and every baseline has p greater than 0.05 in a pairwise T-test, so the paper draws no conclusion about superiority on that metric.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper is about controlled text generation, detoxification, toxicity reduction in language models, sentiment steering, text simplification, contrastive learning, triplet loss, hidden representation transformation, realtoxicityprompts, dexperts, gedi, pplm, gpt-2, inference latency, and multi-attribute control.
What this paper does not show
The abstract says seven baselines; the body compares against five plus base GPT-2. Everything runs on GPT-2 medium, 355M parameters, with nothing on larger or instruction-tuned models, and the attribute scorers are the same classifier families used to build the training data. Human evaluation is a real strength here, and it shows topicality is a statistical tie against all six baselines and that DExperts ties us on linguistic quality. Fluency regresses on sentiment steering. So "outperforms all the baselines" does not hold metric by metric.
How to cite
Vaibhav Kumar, Hana Koorehdavoudi, Masud Moshtaghi, Amita Misra, Ankit Chadha, and Emilio Ferrara. 2023. Controlled Text Generation with Hidden Representation Transformations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9440–9455.
@inproceedings{kumar2023controlled,
title = {Controlled Text Generation with Hidden Representation Transformations},
author = {Kumar, Vaibhav and Koorehdavoudi, Hana and Moshtaghi, Masud and Misra, Amita and Chadha, Ankit and Ferrara, Emilio},
editor = {Rogers, Anna and Boyd-Graber, Jordan and Okazaki, Naoaki},
booktitle = {Findings of the Association for Computational Linguistics: ACL 2023},
month = jul,
year = {2023},
address = {Toronto, Canada},
publisher = {Association for Computational Linguistics},
pages = {9440--9455},
doi = {10.18653/v1/2023.findings-acl.602},
url = {https://aclanthology.org/2023.findings-acl.602/}
}
Related work of mine
- ACM -- Attribute Conditioning for Abstractive Multi Document SummarizationarXiv preprint
- APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AIACL 2026 (Main Conference, Long Papers)
arXiv version. Cited 13 times as counted by Google Scholar on 2026-09-15.