能力Expertise

分两条轴呈现:已验证的能力(有论文/落地产出,证明研究实力)与兴趣与探索方向(我想深耕、也在寻找合作与机会的方向)。等级依据文末的评定标准给出,尽量有据可依。Two axes: track record (published/shipped output that proves capability) and interests I'm pursuing (where I want to go deep and find collaboration). Levels follow the rubric below.

✅ 已验证的能力✅ Track record

生成式 AI · 图像/视频生成Generative AI · Image & Video Generation 精通Proficient

一作提出 LAMIC(AAAI 2026),并在 Onestory 实习中落地进核心产品;目前在可灵从事视频生成方向。微调过 T2I/I2I 生成模型(LoRA / Flux-Kontext / EasyControl)并上线。First-authored LAMIC (AAAI 2026), shipped into Onestory products; now working on video generation at Kling. Fine-tuned and deployed T2I/I2I models (LoRA / Flux-Kontext / EasyControl).

DiffusionDiTT2I / I2IControllable GenerationVideo Generation
扩散模型Diffusion Models 熟练Advanced

在 LAMIC / TAG-WM 中深度使用多模态扩散 Transformer 与扩散反演;熟悉 DDPM/DDIM、Flow Matching 等采样与训练机制。Heavily used multimodal diffusion transformers and diffusion inversion in LAMIC / TAG-WM; familiar with DDPM/DDIM and flow matching.

DDPM / DDIMDiffusion InversionFlow MatchingDiT
AI 安全 · 水印与篡改取证AI Safety · Watermarking & Forensics 精通Proficient

一作提出 TAG-WM(ICCV 2025)与 Flow of Truth(面向图像到视频的时序取证),研究生成图像/视频的水印、版权保护与篡改定位。First-authored TAG-WM (ICCV 2025) and Flow of Truth (temporal forensics for I2V).

Generative WatermarkingTamper LocalizationTemporal Forensics

🧭 兴趣与探索方向🧭 Interests I'm pursuing

我最想深耕、也在寻找合作与机会的方向。Where I want to go deep and am looking for collaboration & opportunities.

多模态生成Multimodal Generation

让 AI 跨模态地生成内容,从当前的图像/视频走向更统一的多模态生成——这是我最想深耕的方向。AI that generates across modalities, moving from image/video toward more unified multimodal generation — the direction I most want to pursue.

Unified GenerationAny-to-AnyCross-Modal

多模态理解Multimodal Understanding

让 AI 跨模态地理解世界;在可灵接触统一「理解-生成」,持续跟进 MLLM / VLM 与视觉 tokenizer。AI that understands the world across modalities; explored unified understanding-and-generation at Kling, tracking MLLM / VLM and visual tokenizers.

MLLMVLMUnified Understanding & Generation

机器人 · 具身智能Robotics · Embodied AI

系统性阅读 VLA、世界-动作模型、模仿学习(占我近期阅读很大比重),正从阅读走向实践。Systematic reading on VLA, world-action models and imitation learning (a large part of my reading base); moving toward practice.

VLAWorld Action ModelsImitation Learning

仿真 · 游戏(可交互世界)Simulation · Games (interactive worlds)

关注可交互世界、游戏环境与 sim-to-real,作为让 AI 生成并栖居于世界的试验场。Interested in interactive worlds, game environments and sim-to-real as testbeds for AI that generates and inhabits worlds.

World ModelsSim-to-RealInteractive Environments

强化学习与决策Reinforcement Learning & Decision Making

关注 RL / RLHF 与决策智能,作为连接「生成」与「交互」的桥梁。Following RL / RLHF and decision making as the bridge between generation and interaction.

RL / RLHFPolicy LearningAgents

📏 评定标准📏 How levels are decided

「已验证的能力」等级并非凭感觉,而是综合:① 论文产出(是否一作/顶会) ② 工程落地(是否上线/产品化) ③ 复现与实践经验 ④ 文献阅读广度。Track-record levels are not vibes — they combine: (1) publications (first-author / top venue), (2) shipped/productionized work, (3) reproduction & hands-on practice, and (4) reading breadth.

精通Proficient

有一作论文 / 落地产出,能独立从想法到实现完成完整研究或工程。First-author papers or shipped systems; can drive a project end-to-end.

熟练Advanced

有较多复现与实践经验,掌握主流方法,能快速上手推进。Solid hands-on and reproduction experience; can ramp up quickly.

进阶中Growing

系统性阅读 + 初步实验/复现,正在积累。Systematic reading plus early experiments; actively building up.

了解Familiar

广泛涉猎,理解核心概念与前沿脉络。Broad exposure; understand core concepts and the research landscape.

🤝 寻求合作与工作机会🤝 Open to collaboration & opportunities

我对多模态生成、机器人/具身、仿真与游戏、强化学习等方向的合作、实习与全职机会持开放态度。欢迎通过 邮件领英 联系我。I'm open to collaboration, internship and full-time opportunities in multimodal generation, robotics/embodied AI, simulation & games, and RL. Reach me via email or LinkedIn.