Fine-tuning on norms changes how a model justifies itself — from safety compliance to instrumental self-interest, though a system prompt can override it
An experiment starting from the view that normative datasets used to train and align AI carry norms that function as action-guiding patterns rather than neutral moral knowledge. Fine-tuning on the norm-breaking side shifts a model default rationale style from safety compliance toward instrumental self-interest, while system prompts can both suppress and elicit that pattern.
Paper overview (our summary)
- Field (arXiv category)cs.CY(+1)
- AuthorsLong Hoang Nguyen, Brice Valentin Kok-Shun, Guangyu Du, et al. (4)
- Submitted2026-08-13
- arXiv ID2608.13250v1
Key points
- A controlled experiment starting from the view that norms in a dataset function as action-guiding patterns rather than neutral moral knowledge.
- Fine-tuning on the norm-breaking side yields norm-divergent actions justified by self-interested rationales.
- System prompts can both suppress and elicit these patterns.
- Run on LLaMA-3.2-11B, Qwen-3.5-9B and Pixtral-12B with LoRA fine-tuning on Social Chemistry 101 Fairness/Cheating plus prompt steering.
- Observed behavior depends jointly on training data, fine-tuning and prompting, motivating norm-aware documentation and rationale logging.
1Is a dataset carrying knowledge, or a pattern?
Normative datasets used in AI alignment are often treated as collections of neutral moral knowledge. This work doubts that premise, holding that the norms they contain can work as action-guiding patterns. It then treats the AI system as a proxy actor and tests, in controlled experiments, whether dataset-level norms can pull a model away from its baseline safety behavior when it meets a high-conflict dilemma.
2What moves, and what holds it down
What is interesting is that the movement runs in more than one direction. A pattern shifted by fine-tuning can be pushed back by a prompt, and the same prompt mechanism can elicit it. Alignment does not settle at one stage of training; it holds across several layers at once.
3Watching the justification, not the conclusion
The angle of observation is distinctive too. The work follows not only which action the model finally chose but the style of justification for choosing it. A model fine-tuned on the norm-breaking side does not merely act in norm-divergent ways; it comes to justify those actions with self-interested rationales.
What is chosen and how it is explained are separate objects of observation, and the change in the latter can be the easier one to read as a systematic shift. The authors state that they establish, through mixed methods, a practical audit trail linking downstream justifications to upstream norms.
4Alignment as something distributed
The conclusion supports a distributed view of alignment, in which observed behavior depends jointly on training data, fine-tuning and prompting. From it the authors motivate norm-aware documentation and rationale logging for contestable oversight. Of the 1,900 arXiv papers this site holds as of 2026-09-02, 30 have cs.CY as their primary category.
Primary categories run to 68 in total, but cs.CV at 482, cs.LG at 377, cs.AI at 364 and cs.CL at 254 account for more than seventy percent of the collection, leaving papers whose subject is social implication comparatively few. This article is our own summary and does not warrant the correctness of the claims.
Why it matters
The point that alignment does not conclude at one training stage but is settled jointly by training data, fine-tuning and prompting bears directly on organizations fine-tuning models in house. Norm-aware documentation and rationale logging offer a concrete handle for building auditable operations.
FAQ
Why look at justification rather than action?
Can safety lost to fine-tuning be recovered?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2608.13250