// LLM Researcher · Tokyo

Chengguang Gan

LLM Researcher @ Techtouch · Ph.D. in Informatics

Research focus Large Language Models Information Extraction Web Agents Mutual Reinforcement Effect

I build and study large language models for information extraction, web agents, and the Mutual Reinforcement Effect — and I ship the datasets and models to back it up.

Chengguang Gan at the Golden Gate Bridge
Tokyo, Japan
// 01

About

I’m an LLM Researcher at Techtouch in Tokyo. I earned my Ph.D. in Informatics in March 2025 from Yokohama National University, advised by Prof. Tatsunori Mori. My work spans natural language processing and large language models — with a focus on information extraction, prompting, multimodal IE, and web agents.

I introduced the Mutual Reinforcement Effect and built the Japanese (and later multilingual) IE Mix datasets — covering sentence, text, sentiment and POS classification, relation and event extraction — then fine-tuned a line of open LLMs on top of them. Everything ships: you can grab the papers, code, datasets, and models on Hugging Face and GitHub.

  • Large Language Models
  • Information Extraction
  • Prompting
  • Mutual Reinforcement Effect
  • Multimodal IE
  • Web Agents
  • Japanese NLP

Reviewer for NeurIPS, COLING, and COLM.

// 02

News

// 03

Publications

Bold = me. Citation counts via Google Scholar (September 9, 2026).

  1. 2026

    How Output Format Confounds Data Quality and Capability in Instruction Tuning

    Chengguang Gan, Hanjun Wei, Yunhao Liang, Qinghao Zhang, Shiwen Ni, Zhixi Cai

    Preprint arXiv
  2. 2026

    Joint Training Is Not Enough: Conditioned Cross-Granularity Training for Multimodal Document Understanding

    Chengguang Gan, Yunhao Liang, Hanjun Wei, Qinghao Zhang, Shiwen Ni

    Preprint arXiv
  3. 2026

    Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem

    Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

    Preprint arXiv
  4. 2026

    Security Tests as Executable Specifications for LLM Code Generation: Benefits, Trade-offs, and Coverage Limits

    Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

    Preprintcited by 1 arXiv
  5. 2026

    Do Code Language Models Follow Tests? Paired Interventions on Program Behavior

    Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni

    Preprint arXiv
  6. 2026
  7. 2026

    MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

    Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni

    Preprint arXiv
  8. 2026

    A Multilingual Dataset and Empirical Validation for the Mutual Reinforcement Effect in Information Extraction

    Chengguang Gan, Sunbowen Lee, Qingyu Yin, Yunhao Liang, Xinyang He, Hanjun Wei, Younghun Lim, Shijian Wang, Hexiang Huang, QingHao Zhang, Shiwen Ni, Tatsunori Mori

    ACL 2026 Findingscited by 2 ACL Anthology
  9. 2026

    GuideWeb: A Benchmark for Automatic In-App Guide Generation on Real-World Web UIs

    Chengguang Gan, Yunhao Liang, Yoshihiro Tsujii, Tatsunori Mori, Shiwen Ni, Hiroki Itoh

    AACL-IJCNLP 2026 Findings (Accepted)cited by 2 arXiv
  10. 2025
  11. 2025
  12. 2025

    RECODE: Leveraging Reliable Self-generated Tests and Fine-Grained Execution Feedback to Enhance LLM-Based Code Generation

    Yunhao Liang, Ruixuan Ying, Takuya Taniguchi, Chengguang Gan, Zhe Cui

    ICIC 2025cited by 8 Springer
  13. 2025

    Exploring Behavior-Driven Development for Code Generation

    Yunhao Liang, Chengguang Gan, Ruixuan Ying, Zhe Cui

    ICIC 2025cited by 5 Springer
  14. 2025

    Retrieval and Distill: A Temporal Data Shift-Free Paradigm for Online Recommendation System

    Lei Zheng, Ning Li, Chengguang Gan, Yong Yu, Weinan Zhang

    APWeb-WAIM 2025cited by 1 Springer
  15. 2025

    M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction

    Chengguang Gan, Zhixi Cai, Yanbin Wei, Yunhao Liang, Shiwen Ni, Tatsunori Mori

    Preprintcited by 1 arXiv
  16. 2025

    Decoding Prokaryotic Whole Genomes with a Product-Contextualized Large Language Model

    Shiwen Ni, Shuaimin Li, Shijian Wang, Wenzhi Xue, Luoxi Zhang, Liyang Fan, Xinping Bi, Yitai Li, Chengguang Gan, Jiarui Jin, Yuan Lu, Ahmadreza Argha, Hamid Alinejad-Rokny, Tong Si, Min Yang, Teng Wang

    bioRxiv bioRxiv
  17. 2024

    Application of LLM Agents in Recruitment: A Novel Framework for Automated Resume Screening

    Chengguang Gan, Qinghao Zhang, Tatsunori Mori

    Journal of Information Processingcited by 173 J-STAGE
  18. 2024

    II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models

    Ziqiang Liu, Feiteng Fang, Xi Feng, Xinrun Du, Chenhao Zhang, Zekun Wang, Yuelin Bai, Qixuan Zhao, Liyang Fan, Chengguang Gan, et al.

    NeurIPS 2024cited by 22 Proceedings
  19. 2024
  20. 2024
  21. 2024

    Demonstrating Mutual Reinforcement Effect through Information Flow

    Chengguang Gan, Xuzheng He, Qinghao Zhang, Tatsunori Mori

    Preprint arXiv
  22. 2023
  23. 2023

    A Few-Shot Approach to Resume Information Extraction via Prompts

    Chengguang Gan, Tatsunori Mori

    NLDB 2023cited by 18 Springer
  24. 2023
  25. 2023
  26. 2023
  27. 2021

    英文履歴書データ抽出システムへの BERT 適用性の検討

    Chengguang Gan, Yoshihide Takahashi

    IPSJ Kansai 2021 IPSJ
// 04

Experience & Education

Experience

  • 2025 — Present
    LLM ResearcherTechtouch
  • 2025
    Data ScientistGeneric Solution
  • 2024 — 2025
    Part-time ResearcherNII Large Language Model Center
  • 2024
    Research Part-timerRIKEN AIP
  • 2022 — 2023
    Research AssistantYokohama National University

Education

  • 2022 — 2025
    Ph.D. in InformaticsGraduate School of Environment and Information Sciences, Yokohama National University
  • 2020 — 2022
    M.S. in Information TechnologyThe Kyoto College of Graduate Studies for Informatics
  • 2014 — 2018
    East China Jiaotong UniversityNanchang, China
// 05

Open Source

Datasets, models, and code behind the papers — free to use.

More on Hugging Face and GitHub.

// 06

Beyond Research

When I’m away from the terminal.

// 07

Get in touch

Open to research collaborations and roles in LLMs, NLP, and agents. Best reached by email — or find me on the platforms below.

ganchengguan@yahoo.co.jp