SKA-VCT: Listening to the Motion
Physical consistency for audio-visual segmentation
A research project on locating sounding objects by aligning audio spectra with motion evidence instead of trusting static visual saliency alone.
AI Explorer and Software Engineering Undergraduate atTongji University.
I work on multimodal intelligence, reasoning behavior, and AI systems for research. I am especially interested in how models connect what they hear, see, and infer, and how those decisions can be made easier to inspect.
My recent work spans audio-visual segmentation, adaptive reasoning budgets, and multi-agent tools for scientific communication. I enjoy moving between research questions and systems that make those questions testable.
Research interests: multimodal learning, LLM reasoning, AI agents, and research tools.

PosterMELD released its paper, code, and evaluation benchmark for editable scientific poster generation.
DraftCode placed third at the AWS Summit Shanghai hackathon and advanced to the Macau round.
Released the reproducible To Think or Not to Think codebase for reasoning-budget studies in Ref-AVS.
Physical consistency for audio-visual segmentation
A research project on locating sounding objects by aligning audio spectra with motion evidence instead of trusting static visual saliency alone.
Pre-decisional reasoning budgets for Ref-AVS
A reproducible Think-Ground-Segment pipeline for comparing zero, short, and long reasoning before grounding referring audio-visual expressions.
Editable, diverse, print-ready paper-to-poster generation
A multi-agent system that turns scientific papers into editable PowerPoint posters while controlling design diversity, content grounding, and print readiness.
I am looking for AI research and engineering internships, especially around multimodal systems, LLM reasoning, agents, and evaluation-focused AI.
Get in touch