Chen Wenjie · 陈文杰
Building interpretable AI systems for proteins, biological data, and computational discovery.为蛋白质和生物数据构建可解释的 AI 系统,让计算真正参与发现。
Questions before models问题先于模型
I work on AI for biology: protein language models, computational protein design, biological representation learning, and trustworthy biomedical AI. The recurring question is simple to state and difficult to answer: what do these models actually know, and how do we prove it?我做生物 AI 方向的研究:蛋白质语言模型、计算蛋白质设计、生物表征学习和可信生物医学 AI。有一个问题反复出现,它很好说清,却很难回答:这些模型究竟知道什么?我们又怎么证明?
I am an undergraduate in Bioengineering at Wuhan Institute of Technology, with an expected graduation date of June 2027. My research path began with SERS/MOF materials experiments and has since moved toward bioinformatics, protein computation, and machine learning.我目前就读于武汉工程大学生物工程专业,预计于 2027 年 6 月本科毕业。我的研究路径始于 SERS/MOF 材料实验,随后逐步转向生物信息学、蛋白质计算与机器学习。
Research experience研究经历
-
SAE-DISTILLSAE-DISTILL
Testing whether sequence-only models can recover selected SAE-derived residue profiles from ESM-C representations.检验仅使用序列的轻量模型能否恢复 ESM-C 表征中选定的 SAE 残基特征谱。
First author — developed the study, implemented the computational workflow, analysed the results, and prepared the manuscript.第一作者——负责研究设计、计算流程实现、结果分析和稿件撰写。
-
PAO1 Anti-virulence DiscoveryPAO1 抗毒力候选物优先级排序
Evidence-stratified computational prioritization of natural-product candidates across virulence-associated PAO1 targets.针对 PAO1 毒力相关靶点,对天然产物候选物进行分层证据支持的计算优先级排序。
First author — developed the multi-criteria prioritization workflow, analysed the computational evidence, and prepared the manuscript.第一作者——构建多指标优先级排序流程,分析计算证据并完成稿件撰写。
-
SERS / MOF materials researchSERS / MOF 材料研究
Experimental research on SERS/MOF composites for food-safety monitoring and small-molecule biomarker detection.参与面向食品安全检测和小分子生物标志物检测的 SERS/MOF 复合材料实验研究。
Co-author — contributed to material preparation, sample testing, data organization, manuscript writing, and supporting electric-field simulations.共同作者——参与材料制备、样品检测、数据整理与论文写作,并辅助完成部分电场模拟。
Working principles工作原则
-
Ask a falsifiable question提出可证伪的问题
Start from what would disprove the idea, not from which model is fashionable.先想清楚什么结果会否定这个想法,而不是先挑当下流行的模型。
-
Make the split resemble use数据划分要贴近真实用法
Family-held-out proteins and external clinical cohorts expose failure modes that random splits can hide.随机划分会掩盖一些失败模式,蛋白质家族留出和外部临床队列能把它们暴露出来。
-
Stress-test the explanation解释也要经受压力测试
Structured controls help separate a recovered biological signal from family texture or benchmark leakage.真正恢复的生物信号、家族统计纹理、基准泄漏,结构化对照能把这三者分开。
-
Stop the claim at the evidence论断止于证据
Computational prioritization, calibration and interpretation each support different kinds of sentences.计算优先排序、校准、解释,各自能撑起的话不一样。
Research interests研究兴趣
- AI for Biology生物人工智能
- Computational Biology计算生物学
- Computational Protein Design计算蛋白质设计
- Protein Language Models蛋白质语言模型
- Biological Representation Learning生物表征学习
- Interpretable Machine Learning可解释机器学习
- Bioinformatics生物信息学
- AI-assisted Drug DiscoveryAI 辅助药物发现