Trustworthy AI
My current focus is the trustworthiness of AI systems, with an interest in understanding their limitations and evaluating the reliability of their outputs.
Current research focus
Evaluating AI. Understanding its limits.
My current research focus is trustworthy AI. I am interested in how we can evaluate AI systems and make their outputs more reliable.
My previous work spans multimodal code understanding, code generation evaluation, and review credibility assessment. These experiences shape my approach to trustworthy AI research.
I am currently an M.Sc. student in Library and Information Studies at Hohai University, and I also work with the LLM for Software Engineering Lab at Shanghai Jiao Tong University.
I study trustworthy AI, building on my work in model evaluation and multimodal learning.
My current focus is the trustworthiness of AI systems, with an interest in understanding their limitations and evaluating the reliability of their outputs.
My work on ClassEval-Pro and CodeOCR examines code generation and multimodal code understanding through benchmarks, test suites, and error analysis.
My work on review credibility combines textual, visual, and relational evidence to assess human-written and AI-generated reviews.
Selected publications and ongoing work.
Proceedings of ISSTA 2026 · 2026
Studies code-as-image representations for multimodal code understanding and shows how visual encoding can improve efficiency while remaining competitive on downstream tasks.
Proceedings of AIware 2026, Benchmark & Dataset Track · 2026
Introduces ClassEval-Pro, a benchmark of 300 class-level code generation tasks across 11 domains, built through an automated three-stage pipeline with complexity enhancement, cross-domain class composition, and real-world GitHub code integration. Each task is validated by an LLM Judge Ensemble and test suites with over 90% line coverage. Experiments on five frontier LLMs under five generation strategies show that the best model reaches only 45.6% class-level Pass@1, while error analysis highlights logic and dependency errors as the main bottlenecks.
International Journal of Intelligent Systems · 2026
Presents MDCFN, a multimodal architecture for robust review credibility assessment across textual, visual, and relational signals.
My research experience and background in Python backend engineering.
LLM for Software Engineering Lab (LLMSE), Shanghai Jiao Tong University
Advisor: Prof. Xiaodong Gu
Institute of Management Science, Hohai University
Hohai University
Hohai University
若界人工智能实验室
Inspur Morning Cloud Technologies Co., Ltd.
Selected work in language models and data analysis.