Experience
I work at the boundary of model research, training systems, and evaluation, with a focus on making AI hold up under real-world constraints.
Professional Work
-
Starting a new chapter in AI research and engineering.
-
LLM compression for on-device deployment
- Compressed an LLM for on-device deployment, combining pruning, distillation and quantization.
- Met the memory target within a tight accuracy budget.
-
VLM compression
- Studied how far 7-8B VLMs compress through quantization, depth pruning and vision-token pruning.
- Depth pruning: VLMs tolerate far fewer removed layers than LLMs. Benchmark scores held while open-ended and reasoning quality fell.
- Vision-token pruning: extended vLLM to run these methods in real serving. Most lost their latency gains to selection overhead and poor batch parallelism.
- The limits of train-free methods became the evidence that shifted the team's direction toward training.
-
Post-training VLMs for industrial domains
- Worked across the full cycle of specializing VLMs with SFT and RL, from training infrastructure and data to evaluation.
- Training infrastructure: established verl as the training stack for better scalability and stronger backends (FSDP, Megatron), and fixed its instabilities, including inference-engine initialization.
- Training data: built the SFT and RL training sets for domain specialization.
- Evaluation: built the evaluation data, a human-labeling loop and an internal benchmark.
- Showed that a MoE-based VLM could curb catastrophic forgetting during domain specialization.
-
Data and evaluation pipeline
- Built an agent-based data pipeline that automates data generation, processing and error analysis.
- Analyzed training data with embedding models, from reasoning-trajectory patterns to vision coverage, to guide synthesis and sampling.
- Takeaway: even at LLM scale, error analysis, human labeling and a well-defined benchmark decide domain specialization.
-
Multimodal protein-ligand modeling
- Modeled proteins and molecules in several forms at once, from strings to 3D atoms and meshes.
- Traced a bias toward one modality through error analysis and reduced it with feature mixing, improving prediction.
-
Target-aware molecule generation
- Explored a diffusion model conditioned on the target protein, to address the existing RL generator's tendency to produce similar molecules.
- Benchmarks were promising, but exposed the gap between academic generation metrics and drug-discovery value.
-
Shared protein-molecule representation
- Learned a joint embedding of small molecules and proteins with contrastive learning and an SO(3)-equivariant design.
- Promising on small-molecule property prediction, while showing that broader gains would need far more data and compute.
Research Foundation
My research foundation grew out of undergraduate work in optimization theory and later graduate research in deep learning theory and generative modeling.
M.S. Research
Sep 2020 - Aug 2022
- Worked on Neural Tangent Kernels and related questions about the training dynamics of overparameterized neural networks.
- Researched generative modeling through diffusion and Schrödinger Bridge methods, culminating in my master's thesis.
B.S. Research
Mar 2016 - Aug 2020
- Explored the theoretical background behind optimizer convergence in deep learning through an undergraduate research program.
- Strengthened my mathematical grounding in machine learning and reinforcement learning.
Teaching
Elementary Mathematical Analysis
Spring 2022
Linear Algebra
Spring 2022, Fall 2021
Supported students through office hours on concepts, proofs, assignments, and exam preparation.
Calculus
Fall 2021, Spring 2021, Fall 2020
Led problem sessions centered on core exercises and problem-solving practice.