I am an assistant professor in the Computer Science Department at Johns Hopkins University, leading MAGIC Lab. I am also affiliated with the Data Science and AI Institute (DSAI), the Laboratory for Computational Sensing and Robotics (LCSR), and the Center for Language and Speech Processing (CLSP).
My research focuses on multimodal AI, integrating diverse data types (e.g., images, videos, text, audio, and motion) to develop models that are interpretable, controllable, and scalable. My recent research interests include:
(1) Scalable Multimodal Frameworks – Modern AI models must meet the growing demand for thousands of capabilities. My research has addressed this challenge by introducing: (a) Unified generative frameworks that flexibly accommodate diverse modalities and tasks, using a single architecture and a generative objective – VL-T5 (ICML 2021) / X-LXMERT (EMNLP 2020) / TVLT (NeurIPS 2022 Oral) and (b) Efficient finetuning frameworks that significantly reduce parameter and memory requirements for creating task-specific models – VL-Adapter (CVPR 2022) / LST (NeurIPS 2022) / Ctrl-Adapter (ICLR 2025 Oral)
(2) Faithful Multimodal Reasoning – Scaling alone is not enough. Large models that rely on black-box reasoning and encode all knowledge within their parameters often struggle with basic tasks and produce hallucinations. My research makes their reasoning process more accurate and interpretable by introducing: (a) Planning-based frameworks that decompose complex visual generation problems into faithful, human-interpretable step-by-step reasoning processes – VPGen (NeurIPS 2023) / VideoDirectorGPT (COLM 2024) / DiagrammerGPT (COLM 2024) / Video-MSG (2024) and (b) Retrieval-augmented generation (RAG) frameworks that enhance accuracy and factuality by retrieving relevant information before generating outputs – M3DocRAG (Findings of ICCV 2025) / HiREST (CVPR 2023)
(3) Evaluation and Refinement of Multimodal Generation – With recent advancements in multimodal generation models, conventional evaluation metrics have been often saturated and no longer provide meaningful insights into future research direction. To this end, my research introduces: (a) Fine-grained evaluation frameworks that comprehensively measure model skills in multiple dimensions to uncover detailed strengths and weaknesses – DALL-Eval (ICCV 2023) / VPEval (NeurIPS 2023) / DSG (ICLR 2024) / LayoutBench (CVPRW 2024 Oral) / FineCapEval (Findings of NAACL 2022) / M3DocVQA (Findings of ICCV 2025) / CAPTURe (ICCV 2025) and (b) Automatic model refinement frameworks that use these evaluations to detect models’ weaknesses and refine their reasoning process – EnvGen (COLM 2024) / DataEnvGym (ICLR 2025 Spotlight) / SELMA (NeurIPS 2024) / VideoRepair (Findings of ACL 2026)
I know these are far from perfect, but I hope this helps you in your applications!
Ph.D. in Computer Science, 2025
University of North Carolina at Chapel Hill
B.S. in Industrial Engineering, 2018
Seoul National University
Young Investigator, 2025 - 2026
Allen Institute for AI (AI2)
Research Intern, 2024
Bloomberg
Student Researcher, 2023
Google Research
Research Intern, 2022
Microsoft Research
Research Intern, 2021
Adobe Research
Predoctoral Young Investigator, 2019 - 2020
Allen Institute for AI (AI2)
Visiting Scholar, 2019
SNU Music & Audio Research Group
AI Resident, 2018 - 2019
Naver Clova
Research Intern, 2017 - 2018
SNU Vision & Learning Lab
Sep 2026 - 2 papers accepted at NeurIPS 2026:
Sep 2026 - 2 papers accepted at CoRL 2026:
Jul 2026 - Jaemin joined the Computer Science Department at Johns Hopkins University as an assistant professor.
Jun 2026 - 4 papers accepted at ECCV 2026:
Jun 2026 - New preprint:
May 2026 - New preprints:
May 2026 - 1 paper accepted at ICML 2026:
Apr 2026 - New preprint:
Apr 2026 - 1 paper accepted at Findings of ACL 2026:
Apr 2026 - 1 paper accepted at CVPR 2026 MUSI Workshop:
Mar 2026 - Jaemin received the UNC Dean’s Distinguished Dissertation Award for his dissertation, “Modular and Interpretable Multimodal AI for Improved Generation and Evaluation.”
Mar 2026 - New preprints:
Feb 2026 - New preprint:
Jan 2026 - 1 paper accepted at ICLR 2026:
Jan 2026 - 1 paper accepted at EACL 2026:
Dec 2025 - Panel discussion at DCVLR: Data Curation for Vision Language Reasoning NeurIPS 2025 Workshop
Sep 2025 - 1 paper accepted at NeurIPS 2025:
Sep 2025 - Starting gap year at AI2 PRIOR team as a Young Investigator!
Jun 2025 - New preprint:
Feb 2025 - 2 papers accepted at ICLR 2025:
MAGIC (Multimodal AI for Grounded Intelligence & Creation) Lab is a research group at Johns Hopkins University led by Prof. Jaemin Cho.
PhD Students
Zhanpeng Luo (Fall 2026)
Lin Long (Fall 2026)
Regarding joining the group, please see this page.
2026, Univ. of Delaware, “Modular and Interpretable Multimodal AI for Improved Generation and Evaluation”
2026, Keynote at ECCV Workshop on Multimodal LLMs for Unified Comprehension and Generation, “Unify the Interface, Not Just the Parameters: A Shared Symbolic Canvas for Human-AI Visual Co-Creation”
2026, Keynote at ECCV Workshop on Curated Data for Efficient Learning, “Fine-Grained Evaluation and Targeted Curriculum for Efficient Learning”
2026, Naver, “Toward the GPT Moment of Physical AI”
2025, Panel discussion at NeurIPS Workshop on Data Curation for Vision Language Reasoning
2025, GNU, CMU, UW, Emory, SKKU, SNU, Korea Univ., Yonsei Univ., Krafton, UIUC, UNC, JHU, RPI, Purdue, “Modular and Interpretable Multimodal AI for Improved Generation and Evaluation”
2024, Korea Univ., “Language Model as Your Multimodal Generator, Evaluator, and Trainer”
2023, UNC, Korea Univ., “Hierarchical Video-Moment Retrieval and Step-Captioning”
2022, POSTECH, Chung-Ang Univ., Naver, KAIST, Korea Univ., “Vision-and-Language: Pretraining, Transfer Learning, and Evaluation”
2021, Kakao, “Unifying Vision-and-Language Tasks via Text Generation”
2020, UNC, “X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers”
2020, SNU, UNIST, “Mixture Content Selection for Diverse Sequence Generation”
2018, Kakao, “Generative Models & Variational Inference”
2017, TensorFlow KR, “Developing Korean Chatbot 101”
Please see Google Scholar or Semantic Scholar for an up-to-date list.