AC

Controlled Text Generation with Hidden Representation Transformations

Controlled generationFindings of ACL 2023, pages 9440–9455 · Amazon Alexa AI; University of Southern California

PDF arXiv DOI Publisher page Code BibTeX

In one sentence

CHRT steers a language model away from toxic or negative text by editing its internal number vectors, cutting toxicity from 0.827 to 0.085 while adding only 0.01 seconds per generation.

Abstract

We propose CHRT (Control Hidden Representation Transformation) – a controlled language generation framework that steers large language models to generate text pertaining to certain attributes (such as toxicity). CHRT gains attribute control by modifying the hidden representation of the base model through learned transformations. We employ a contrastive-learning framework to learn these transformations that can be combined to gain multi-attribute control. The effectiveness of CHRT is experimentally shown by comparing it with seven baselines over three attributes. CHRT outperforms all the baselines in the task of detoxification, positive sentiment steering, and text simplification while minimizing the loss in linguistic qualities. Further, our approach has the lowest inference latency of only 0.01 seconds more than the base model, making it the most suitable for high-performance production environments. We open-source our code and release two novel datasets to further propel controlled language generation research.

Why it matters

Most ways to stop a language model from producing toxic text either slow generation down a lot or make the text read worse. CHRT trains a small block that rewrites the model's internal representations, so the base model and its decoder stay untouched at run time. On the paper's tests it cut toxicity further than five prior methods while adding 0.01 seconds per continuation, against 9.30 extra seconds for PPLM. Attribute blocks can also be mixed at inference time without retraining.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

CHRT freezes the base model and trains only a transform block on its hidden states, guided by two fine-tuned models.
Figure 1: Visual Representation for CHRT’s Training, Inference and Transformation Block. Taken from the paper.

Topics

This paper is about controlled text generation, detoxification, toxicity reduction in language models, sentiment steering, text simplification, contrastive learning, triplet loss, hidden representation transformation, realtoxicityprompts, dexperts, gedi, pplm, gpt-2, inference latency, and multi-attribute control.

What this paper does not show

The abstract says seven baselines; the body compares against five plus base GPT-2. Everything runs on GPT-2 medium, 355M parameters, with nothing on larger or instruction-tuned models, and the attribute scorers are the same classifier families used to build the training data. Human evaluation is a real strength here, and it shows topicality is a statistical tie against all six baselines and that DExperts ties us on linguistic quality. Fluency regresses on sentiment steering. So "outperforms all the baselines" does not hold metric by metric.

How to cite

Vaibhav Kumar, Hana Koorehdavoudi, Masud Moshtaghi, Amita Misra, Ankit Chadha, and Emilio Ferrara. 2023. Controlled Text Generation with Hidden Representation Transformations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9440–9455.

@inproceedings{kumar2023controlled,
  title     = {Controlled Text Generation with Hidden Representation Transformations},
  author    = {Kumar, Vaibhav and Koorehdavoudi, Hana and Moshtaghi, Masud and Misra, Amita and Chadha, Ankit and Ferrara, Emilio},
  editor    = {Rogers, Anna and Boyd-Graber, Jordan and Okazaki, Naoaki},
  booktitle = {Findings of the Association for Computational Linguistics: ACL 2023},
  month     = jul,
  year      = {2023},
  address   = {Toronto, Canada},
  publisher = {Association for Computational Linguistics},
  pages     = {9440--9455},
  doi       = {10.18653/v1/2023.findings-acl.602},
  url       = {https://aclanthology.org/2023.findings-acl.602/}
}

Related work of mine

arXiv version. Cited 13 times as counted by Google Scholar on 2026-09-15.

← All publications