Research研究方向

Questions I work on我关注的问题

Four directions, one thread: how computational models represent, predict and design biological systems — and how to prove it rigorously.四个方向,一条主线:计算模型如何表征、预测和设计生物系统,以及怎么把这件事证明得足够严格。

Protein Language Models & Representation Learning蛋白质语言模型与表征学习

The question研究问题

What do protein language models actually encode about biology, and how can we recover it faithfully?蛋白质语言模型究竟编码了哪些生物信息?我们如何可靠地恢复并验证这些信息?

Why it matters为什么重要

PLMs now underpin structure prediction, variant effect scoring, and design. If we cannot say precisely what they encode — and verify it under strict controls — downstream conclusions rest on unexamined representations.PLM 已用于结构预测、变异效应评分和蛋白质设计。如果无法准确说明并严格验证它们编码了什么,下游结论就会建立在未经检验的表征之上。

My focus我的关注点

  • Residue-level representation recovery残基级表征恢复
  • Sparse autoencoder feature analysis稀疏自编码器特征分析
  • Structured-null and family-held-out evaluation结构化零假设与家族留出评估
  • Cross-family generalization跨家族泛化

I study how sequence-trained models represent biological signals at residue resolution, and how to distinguish genuine encoding from evaluation artifacts. This means building probes that must survive structured nulls and family-held-out splits before any biological claim is made.

Current work centers on SAE-DISTILL, where sparse autoencoder features are used as the unit of analysis for recoverability questions.

我研究只用序列训练的模型如何在残基层面表示生物信号,以及如何区分真实编码与评估伪影。这意味着,任何生物学论断提出之前,探针都必须先通过结构化零假设和家族留出划分的检验。

目前的工作以 SAE-DISTILL 为主,使用稀疏自编码器特征作为分析单元,研究不同信号的可恢复性。

Conceptual sequence field transforming into a sparse feature lattice
Sequence → representation → interpretable feature序列 → 表征 → 可解释特征 AI-generated conceptual visual · not experimental dataAI 生成概念视觉 · 非实验数据

Computational Protein Design计算蛋白质设计

The question研究问题

How do we move from predicting protein structure to designing sequences and structures with intended function?如何从预测蛋白质结构,进一步走向设计具有目标功能的序列和结构?

Why it matters为什么重要

Design closes the loop of the sequence–structure–function paradigm. Reliable computational design would turn biology from an observational science into an engineering discipline.设计把序列—结构—功能关系连成了闭环。可靠的计算设计有望让生物学从观察性科学进一步走向可控设计。

My focus我的关注点

  • Structure generation and sequence design结构生成与序列设计
  • Design validation via structure prediction利用结构预测验证设计结果
  • Sequence–structure co-consistency checks序列—结构一致性检查

My interest here is the full design loop: generating backbones, designing sequences onto them, and validating designs by re-prediction. I follow diffusion-based generators and inverse folding models closely, with particular attention to how design success is actually validated — self-consistency metrics versus independent predictors.

我关注完整的蛋白质设计流程:生成骨架、为骨架设计序列,再通过重新预测结构来验证设计。我也持续关注基于扩散的生成模型和逆折叠模型,尤其重视如何验证所谓的设计成功,例如区分自洽性指标与独立预测器给出的证据。

Conceptual protein ribbon emerging from geometric constraints
Constraint → generation → independent validation约束 → 生成 → 独立验证 AI-generated conceptual visual · not experimental dataAI 生成概念视觉 · 非实验数据

Bioinformatics & Biological Data生物信息学与生物数据

The question研究问题

How can sequence, structure, and molecular data be curated into analyses with clear evidence boundaries?如何把序列、结构和分子数据整理成证据边界清晰的分析?

Why it matters为什么重要

Biological datasets are noisy, redundant, and confounded. Careful curation and honest evaluation design are essential when moving from a computational score to a scientific claim.生物数据集往往噪声大、冗余高,而且存在复杂混杂因素。从计算分数走向科学论断时,谨慎的数据整理和诚实的评估设计不可缺少。

My focus我的关注点

  • Protein sequence analysis蛋白质序列分析
  • Structural and molecular computation结构与分子计算
  • Biological data curation and benchmarking生物数据整理与基准构建

My current work connects protein sequences, model-derived residue profiles, molecular structures, and natural-product records. The emphasis is on traceable curation, matched controls, reproducible computational workflows, and conclusions that stop at the available evidence.

我目前处理的数据包括蛋白质序列、模型导出的残基特征谱、分子结构和天然产物记录。工作重点是可追溯的数据整理、匹配对照、可复现计算流程,以及让结论止于现有证据。

Conceptual network of omics matrices, sequences and biological data layers
Heterogeneous data → deployment-shaped benchmark异质数据 → 贴近部署的基准 AI-generated conceptual visual · not experimental dataAI 生成概念视觉 · 非实验数据

Biomedical AI & Trustworthy Validation生物医学 AI 与可信验证

The question研究问题

What makes an AI result in biomedicine convincing enough to act on?什么样的生物医学 AI 结果,才足以支持进一步行动?

Why it matters为什么重要

Computational biomedical results can fail through dataset shift, miscalibration, and benchmark leakage. Validation design determines what evidence a model or ranking can support.生物医学计算结果可能受到数据分布变化、校准不足和基准泄漏影响。评估设计决定了模型或排序结果能够支持什么层级的证据。

My focus我的关注点

  • Knowledge-distillation evaluation design知识蒸馏评估设计
  • External validation and calibration as planned evidence作为计划证据的外部验证与校准
  • Computational drug discovery with explicit claim boundaries带明确论断边界的计算药物发现

This direction includes a method concept for evaluating pathology-model distillation and a completed computational prioritization workflow for PAO1 anti-virulence candidates. Their evidence levels differ: the first specifies tests that future work should run, while the second ranks hypotheses for experimental follow-up without claiming biological validation.

这一方向包括一项病理模型蒸馏的评估方法构想,以及一项已经完成计算流程的 PAO1 抗毒力候选物优先级排序研究。两者的证据阶段不同:前者说明未来应开展哪些检验,后者为实验验证排列假设优先级,但不声称已经得到生物学验证。

Conceptual evidence boundary formed by calibration curves and external cohorts
Prediction → stress test → bounded claim预测 → 压力测试 → 有边界的论断 AI-generated conceptual visual · not experimental dataAI 生成概念视觉 · 非实验数据

Evaluation philosophy评估理念

Evidence I trust我信任的证据

Five checks that keep an appealing result from becoming an oversized claim.五项检查,避免一个漂亮结果被写成超出证据的论断。

  1. 01

    Cross-family generalization跨家族泛化

    For protein models, the split should test transfer beyond familiar sequence families, not memorization within them.蛋白质模型的数据划分要检验它能否迁移到陌生的序列家族,而不是它在熟悉家族里记住了什么。

  2. 02

    External validation外部验证

    For biomedical models, evidence should survive a cohort or setting that did not shape model development.生物医学模型的证据,要能在没参与过模型开发的队列或场景里站得住。

  3. 03

    Calibration校准

    A useful probability must describe uncertainty honestly; discrimination alone is not enough.概率要有用,就得诚实地表达不确定性。只有区分能力是不够的。

  4. 04

    Reproducibility可复现性

    The path from data and controls to the reported claim should be inspectable, including negative outcomes.从数据和对照到最终论断,整条路径都要能查,阴性结果也不例外。

  5. 05

    Claim boundaries论断边界

    The strength of the sentence should never exceed the strength of the experiment that supports it.话说到多重,取决于实验撑得起多重。