Haojie Hu

AI Explorer and Software Engineering Undergraduate atTongji University.

I work on multimodal intelligence, reasoning behavior, and AI systems for research. I am especially interested in how models connect what they hear, see, and infer, and how those decisions can be made easier to inspect.

My recent work spans audio-visual segmentation, adaptive reasoning budgets, and multi-agent tools for scientific communication. I enjoy moving between research questions and systems that make those questions testable.

Research interests: multimodal learning, LLM reasoning, AI agents, and research tools.

Haojie Hu's GitHub avatar
Undergraduate researcher · AI builder

news

  1. PosterMELD released its paper, code, and evaluation benchmark for editable scientific poster generation.

  2. DraftCode placed third at the AWS Summit Shanghai hackathon and advanced to the Macau round.

  3. Released the reproducible To Think or Not to Think codebase for reasoning-budget studies in Ref-AVS.

selected work

All projects
Research project2026

SKA-VCT: Listening to the Motion

Physical consistency for audio-visual segmentation

A research project on locating sounding objects by aligning audio spectra with motion evidence instead of trusting static visual saliency alone.

  • Audio-Visual Segmentation
  • Motion
  • Multimodal Learning
Research codebase2026

To Think or Not to Think

Pre-decisional reasoning budgets for Ref-AVS

A reproducible Think-Ground-Segment pipeline for comparing zero, short, and long reasoning before grounding referring audio-visual expressions.

  • LLM Reasoning
  • Interpretability
  • Ref-AVS
Research system2026

PosterMELD

Editable, diverse, print-ready paper-to-poster generation

A multi-agent system that turns scientific papers into editable PowerPoint posters while controlling design diversity, content grounding, and print readiness.

  • Multi-Agent Systems
  • Scientific Communication
  • PPTX

currently

I am looking for AI research and engineering internships, especially around multimodal systems, LLM reasoning, agents, and evaluation-focused AI.

Get in touch