Research series
Training Data Attribution for NLP and LLM Research
A thesis- and interview-oriented path through attribution units, utilities, Shapley values, ablation, causal evidence, instance-level methods, scaling, uncertainty, validation, and research extensions.
Part I
Define the attribution problem
Part II
Contribution under coalition context
Part III
Methods, scaling, and uncertainty
Part IV
Validation and interview narrative
How to use this series
The goal is not to sound like a generic tutorial. The goal is to make my thesis explainable, defensible, and expandable.
Interview readiness
Each note ends with a spoken-style answer that can be adapted for project interviews and PhD discussions.
Research precision
The series separates attribution units, utilities, estimators, interventions, and uncertainty so the claims stay precise.
PhD direction
The later notes point toward hierarchical attribution, intervention validation, and LLM factuality/style attribution.
Start by defining what attribution is allowed to claim.
A strong attribution project begins with clear units, clear utilities, and careful claims.