We look inside open-weight models, change what we find, and measure the consequences.
Contrastive texts based on Schwartz values and Moral Foundations.
Linear probes locate where values are encoded in each model.
Value vectors injected into activations shift its behaviour.
Human annotators and LLMs compared on "gold standard" datasets.
Steered models tested on classification, polarisation and hate speech.
Annotation experiments (with Platoniq Foundation) and survey experiments on how people perceive biased chatbots. Approved by the University of Barcelona ethics board; all data anonymised and GDPR-compliant.