Yifan Le
Recent Focus
I focus on neuron-level analysis for large language models, exploring how internal mechanisms relate to multilingual ability, mathematical reasoning, structured generation, and post-training behavior.
My recent work studies how relevance-based neuron discovery, causal intervention, and internal activation signals can help explain and improve model behavior after post-training.
Personal Background
I have a research and engineering background in natural language processing and large language models. My current interests connect interpretability, post-training, multilingual modeling, reasoning alignment, and structured generation.
- Zhejiang University, M.S. in Computer Technology (2018.09 – 2021.03)
- Northwestern Polytechnical University, B.Eng. in Mechatronic Engineering (2012.09 – 2016.06)
Research & Papers
YFPO: Neuron-Guided Preference Optimization for Mathematical Reasoning [Working Paper]
Proposes a neuron-guided preference optimization framework for mathematical reasoning. The work uses AttnLRP to identify math-related neurons and introduces the chosen/rejected activation gap as an auxiliary reward signal, connecting external preference data with internal capability representations.
CRANE: Causal Relevance Analysis of Language-Specific Neurons in Multilingual Large Language Models [Preprint]
Studies language-specific neurons in multilingual LLMs through relevance analysis and masking interventions. The project reframes language specificity from activation correlation to functional necessity and introduces LangSpec-F1 to measure target-language degradation and non-target stability.
Schema Key Wording as an Instruction Channel in Structured Generation under Constrained Decoding [Preprint]
Investigates how schema key wording affects structured generation under constrained decoding. The work treats schema keys as an implicit instruction channel and analyzes interactions among prompt instructions, schema-key instructions, reasoning quality, and format adherence.
COMPASS-V2 Technical Report [Technical Report]
Contributed to technical report content around post-training framework optimization, multilingual capability improvement, evaluation loops, and MOA-style post-processing practices.
Competition Experience
Competition experience is listed as historical background and evidence of practical modeling ability. My current focus is on research and writing around neuron-level analysis for LLMs.
Kaggle Competitions Master
Long-term competition experience in NLP, question answering, sequence labeling, weak supervision, data augmentation, model ensembling, and metric-driven iteration.
2023 Global Intelligent Vehicle AI Challenge
Built a coarse-to-fine RAG system for automotive manual QA, combining span-level augmentation, prompt optimization, embedding fusion, reranking, decoding optimization, and metric-aware iteration.
Kaggle Feedback Prize - Predicting Effective Arguments
Designed span-aware NLP pipelines for argument effectiveness classification, including merge-span modeling, prompt-like templates, back translation, model stacking, and robust generalization strategies.
Kaggle Feedback Prize - Evaluating Student Writing
Modeled discourse element extraction as a sequence labeling task, using NER-style modeling, data augmentation, and multi-stage training to improve robustness.
Kaggle chaii - Hindi and Tamil Question Answering
Handled multilingual QA with weak supervision, MML / HardEM-style label-noise reduction, cross-lingual transfer, back translation, language-specific tokenization analysis, and post-processing.
Kaggle Coleridge Initiative - Show US the Data
Worked on dataset mention extraction and retrieval-oriented NLP modeling, strengthening experience in information extraction, noisy-text matching, and document-level evidence modeling.
2021 Future Cup AI Academic League
Achieved first prize in an academic AI competition, reflecting strong practical ability in model development, experiment iteration, and task-specific optimization.
CCKS 2021 Chinese Medical Popular Science Reading Comprehension
Worked on Chinese medical reading comprehension, involving domain-specific QA modeling, evidence matching, and robust answer extraction.
CCKS 2020 Experimental Identification NER
Built named entity recognition systems for specialized text, focusing on span extraction, sequence labeling, domain adaptation, and evaluation-driven optimization.
Beijing Digital Medical Insurance Innovation Competition
Participated in medical-insurance-related NLP / AI modeling, involving domain understanding, data processing, and practical system-oriented optimization.
Epidemic Government Affairs QA Assistant
Worked on question answering for government-affairs scenarios, requiring retrieval, semantic matching, and robust response generation under domain-specific constraints.
iFLYTEK Big Data Application Classification Annotation Challenge
Participated in classification and annotation modeling, focusing on data analysis, feature construction, and reliable validation.
TensorFlow 2.0 Question Answering
Worked on question answering with transformer-based models and competition-style validation, strengthening experience in QA modeling and error analysis.
Professional Background
This section is intentionally brief. The homepage is organized around research interests, competition experience, and public research output rather than current employment.
- Shopee — LLM post-training, multilingual capability improvement, SFT/RLHF/DPO data optimization, evaluation feedback loops, and Megatron-based training optimization.
- Shanghai AI Lab — General LLM fine-tuning, RLHF workflow, CodeLLM data construction, RAG-finetuning, evaluation, and vLLM-based deployment.
- NetEase — NLP generation, intent classification, named entity recognition, customer-service QA, and production-oriented NLP systems.
Site Statistics
Page views and recent visits are tracked privately in GoatCounter.