Data Scaling for LoRA Domain Adaptation
A research note on how to study the effect of domain data size in LoRA-based machine translation adaptation.
Foundations + Research
A bilingual knowledge base on mathematics, statistics, machine learning, NLP, LLMs, model evaluation, and applied AI research. The goal is simple: make the foundations reusable, then connect them to real experiments.
Foundations
This track is where existing high-quality resources become my own explanations, notes, examples, and interview-ready mental models.
Linear algebra, calculus, optimization, graphs, and representation spaces explained through ML and NLP examples.
Probability, distributions, confidence intervals, bootstrap testing, hypothesis tests, and uncertainty in evaluation.
Splits, generalization, model selection, cross-validation, losses, regularization, and practical training workflows.
Tokenization, embeddings, Transformers, fine-tuning, prompting, sequence models, and language-specific evaluation.
Experiment repositories, config-driven pipelines, logging, annotation sheets, demos, and reproducible analysis.
Research & Applications
This track avoids boxing the blog into one thesis topic. It groups research by capability: building, evaluating, explaining, retrieving, and applying AI systems.
Machine translation, multilingual NLP, low-resource adaptation, terminology, domain adaptation, and language variation.
Metrics, benchmark design, statistical testing, ablations, error analysis, human evaluation, and robustness.
Interpretability, bias analysis, attribution, model auditing, SHAP/LIME, Shapley values, and controlled evaluation.
Information retrieval, RAG, long documents, graph retrieval, evidence grounding, and document intelligence.
Data pipelines, classification, forecasting, information extraction, deployment, product demos, and applied analytics.
Posts may be in English, Chinese, or bilingual. Foundations posts often start from strong public resources; research posts connect those foundations to my own experiments.
A research note on how to study the effect of domain data size in LoRA-based machine translation adaptation.
A research-story version of my LoRA NMT project: problem framing, data, method, evaluation, findings, limitations, and future work.
Why terminology should be evaluated explicitly in domain machine translation, especially when automatic metrics can hide critical technical errors.
A research-facing explanation of why LoRA is a good fit for low-resource domain machine translation: controlled adaptation, lower experimental cost, and reduced overfitting risk.
如何研究领域数据规模对 LoRA 机器翻译适配的影响:数据多少才够,什么时候收益递减,什么时候是数据质量问题。
用研究叙事方式解释我的 LoRA NMT 项目:问题定义、数据、方法、评估、发现、局限与未来工作。
Portfolio pages show what I built. These notes explain the foundations, decisions, experiments, and tradeoffs behind the work.