Methods

How we study it

We look inside open-weight models, change what we find, and measure the consequences.

01

Five steps

  1. 01

    Build a value corpus

    Contrastive texts based on Schwartz values and Moral Foundations.

    WP2
  2. 02

    Probe the layers

    Linear probes locate where values are encoded in each model.

    WP3
  3. 03

    Steer the model

    Value vectors injected into activations shift its behaviour.

    WP4
  4. 04

    Audit benchmarks

    Human annotators and LLMs compared on "gold standard" datasets.

    WP5
  5. 05

    Measure the impact

    Steered models tested on classification, polarisation and hate speech.

    WP6
02

Models & infrastructure

ModelsLlama · Mistral · Gemma · Qwen
Scale7B to hundreds of billions
ComputeMareNostrum 5, BSC
AccessOpen weights, run locally
03

Human studies & ethics

Annotation experiments (with Platoniq Foundation) and survey experiments on how people perceive biased chatbots. Approved by the University of Barcelona ethics board; all data anonymised and GDPR-compliant.