Hi! π
I am a fourth year computer science PhD student at Vislang Lab at Rice University, advised by Prof. Vicente Ordonez.
My research focuses on self-improving AI agents and multimodal reasoning, with an emphasis on hypothesis-driven discovery, tool use, and reliable evaluation. I am particularly working on:
π Self-Improving and Discovery-Driven Agents
- Discover neural architectures through hypothesis generation, implementation, and experimental feedback [HypoExplore].
- Automate hierarchical music mixing through audio-processing tool planning and parameter refinement with a learned critic (Sony AI; details coming soon).
π§ Multimodal Reasoning and Tool Use
- Generate property tests to guide visual program generation and detect errors in tool-based reasoning [PropTest].
- Improve reasoning through reinforcement learning with softmax-based advantage estimation [SoftmaxGRPO] and inference-time guidance from small visual reasoners without additional training [ProxyThinker].
- Ground objects through scenario understanding [RSC], extract event hierarchies across text and video [Beyond Grounding], and fuse audio and text for sentiment analysis [MMML].
π Reliable Evaluation of Generative Models and Agents
- Evaluate visual fidelity and prompt alignment jointly using conditional FrΓ©chet distance, with strong agreement with human judgments on image and video generation benchmarks [cFreD].
- Develop an LLM-as-a-judge method for evaluating music post-production quality, where there is no single correct mix (Sony AI; details coming soon).
I completed my MS in Computer Science from Columbia University, where I worked on multimodal emotion detection in Speech Lab, advised by Prof. Julia Hirschberg. I have also worked with Prof. Shih-Fu Chang at Columbia University.
Previously, I graduated from Ewha Womans University with B.S. in Computer Science and Engineering, and Scranton Honors Program , where I was advised by Prof. Dongbo Min and by Prof. Hyun-Seok Park .
π₯ News
- [08/2026] Our paper HypoExplore is accepted to EMNLP 2026 Findings and "Beyond Referring Expressions: Scenario Comprehension Visual Grounding" is accepted to EMNLP 2026 Main. π
- [07/2026] Our paper "SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation" is accepted to COLM 2026. π
- [05/2026] I joined SONY AI as a research intern this summer.
- [01/2026] Our paper "ProxyThinker: Test-Time Guidance through Small Visual Reasoners" is accepted to ICLR 2026. π
- [09/2025] Our paper "Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Frechet Distance" is accepted to WACV 2026. π
- [09/2024] Our paper "PropTest: Automatic Property Testing for Improved Visual Programming" is accepted to EMNLP 2024 Findings. π
- [03/2024] Our paper "Multi-Modality Multi-Loss Fusion Network" is accepted to NAACL 2024 (Oral). π
π Publications
Preprints
Publications


Ruozhen He, Nisarg A Shah, Qihua Dong, Zilin Xiao, Jaywon Koo, Vicente Ordonez
The 2026 Conference on Empirical Methods in Natural Language Processing. EMNLP 2026.
Paper | Project Page

Jefferson Hernandez, Jaywon Koo, Zilin Xiao, Chen Wei, Vicente Ordonez
Conference on Language Modeling. COLM 2026.
Paper



Jaywon Koo, Ziyan Yang, Paola Cascante-Bonilla, Baishakhi Ray, Vicente Ordonez
Findings of Empirical Methods in Natural Language Processing. Findings of EMNLP 2024.
Paper | Project Page


Hammad Ayyubi, Christopher Thomas, Lovish Chum, Rahul Lokesh, Long Chen, Yulei Niu, Xudong Lin, Xuande Feng, Jaywon Koo, Sounak Ray, Shih-Fu Chang
Proceedings of the AAAI Conference on Artificial Intelligence. AAAI-24.
Paper
