· Xiaojing Yang · Explainability and Responsible AI · 4 min read

中文

Attribution Units: From Data Sources to Individual Examples

Why attribution units matter: source, group, document, example, and token-level attribution answer different research questions and support different kinds of evidence.

Core idea

The attribution unit is not a technical detail. It defines the scientific claim. A source-level claim, a document-level claim, and an example-level claim are not interchangeable.

This note is part of my series Training Data Attribution for NLP and LLM Research. The series is written as both a research notebook and an interview preparation path: each article should help me explain the idea clearly, connect it to my thesis, and identify what would become future PhD work.

Guiding question: What exactly receives credit or blame in training-data attribution?

Training data attribution map

Intuition

If I say a group of Norwegian petroleum data helped domain terminology, I am making a different claim from saying one sentence pair caused one translation. Coarser units are often more stable and interpretable, while finer units can be more actionable but noisier and harder to validate.

In NLP and LLM research, this matters because model behaviour is deeply shaped by data mixture. A model may be fluent because of broad web text, domain-accurate because of specialised documents, safer because of curated instruction data, or biased because of repeated patterns in a subset of the corpus. Training-data attribution gives us language for asking these questions systematically instead of only saying “the data matters”.

Formal lens

Let the training set be partitioned into units G = {g_1, …, g_m}. Each g_i may be a source, group, document, example, or token subset. Attribution estimates phi_i for each unit. Changing the partition changes the game: splitting one source into documents can reveal heterogeneity, but it can also increase variance and computational cost.

The important discipline is to define the attribution setup before interpreting the score:

Design choiceQuestion to answer
Attribution unitWhat receives credit: source, group, document, example, or token?
Utility functionWhich behaviour is being explained: quality, terminology, style, factuality, or safety?
InterventionAre we adding, deleting, reweighting, correcting, or retraining?
EstimatorIs the score exact, sampled, gradient-based, surrogate-based, or heuristic?
UncertaintyHow stable is the score across seeds, samples, metrics, and evaluation sets?

NLP / LLM example

For a domain MT thesis, a practical hierarchy is source -> document family -> document -> sentence pair -> term/token. The top level helps answer research questions about data collection. The middle level helps audit corpora. The bottom level helps inspect examples that may explain a particular error.

This is why I do not want to treat attribution as a generic interpretability topic. For my profile, the natural connection is multilingual and domain-specific NLP: low-resource settings, technical terminology, written-standard variation, and evaluation beyond one headline metric.

Connection to my thesis

In my thesis narrative, training-data attribution is useful because it turns a vague data question into an experimental design:

  1. define interpretable data units;
  2. define the model behaviour to explain;
  3. compare controlled data coalitions or interventions;
  4. estimate contribution;
  5. report uncertainty and limitations;
  6. decide what evidence is strong enough to support a causal-style claim.

That structure helps me avoid overclaiming. A score is not automatically a causal explanation. It is a measurement produced by a specific setup.

What I have done, understand, and would extend

LevelStatus
Already completed / thesis-readyGroup-level attribution, coalition thinking, metric-based utilities, cautious interpretation, random baselines, bootstrap-style reliability checks.
I understand but may not fully implement yetInstance-level gradient attribution, influence functions, TracIn, Monte Carlo Shapley, surrogate/datamodel approximations.
Strong PhD extensionHierarchical attribution, intervention-based validation, factuality/style-specific utilities, scalable attribution for LLM data mixtures.

Interview answer

I would explain my attribution unit before reporting any score. In my thesis setting, group-level units are appropriate because I care about interpretable data sources and written-standard groups, not only isolated training examples. Instance-level methods are useful later for diagnosis, but they should be nested inside a broader group-level picture.

References and reading path

  • Lloyd Shapley, A Value for n-Person Games.
  • Ghorbani and Zou, Data Shapley: Equitable Valuation of Data for Machine Learning.
  • Koh and Liang, Understanding Black-box Predictions via Influence Functions.
  • Pruthi et al., Estimating Training Data Influence by Tracing Gradient Descent.
  • Ilyas et al., Datamodels: Predicting Predictions from Training Data.
  • Rei et al., COMET: A Neural Framework for MT Evaluation.
Share:
Back to Blog

Related Posts

View All Posts »
Research & ApplicationsExplainability and Responsible AIEN

Coalitions, Marginal Contributions, and Shapley Values

The basic Shapley framework for data attribution: coalition value, marginal contribution, averaging over contexts, and why the result is more stable than one ablation.

Research & ApplicationsExplainability and Responsible AIEN

Data Intervention and Attribution Validation

How to validate data attribution through deletion, correction, reweighting, counterfactual examples, retraining, and random deletion baselines.