Chaoxiang Xie

Current research focus

Trustworthy AI

Evaluating AI. Understanding its limits.

My current research focus is trustworthy AI. I am interested in how we can evaluate AI systems and make their outputs more reliable.

My previous work spans multimodal code understanding, code generation evaluation, and review credibility assessment. These experiences shape my approach to trustworthy AI research.

  • Trustworthy AI
  • Model Evaluation
  • Multimodal Learning

I am currently an M.Sc. student in Library and Information Studies at Hohai University, and I also work with the LLM for Software Engineering Lab at Shanghai Jiao Tong University.

News

  • 2026-08 Joined 若界人工智能实验室 as an Agent Algorithm Intern.
  • 2026-04-21 ClassEval-Pro was accepted to AIWare 2026.
  • 2026-04-17 CodeOCR was accepted to ISSTA 2026.
  • 2025-10 Joined the LLM for Software Engineering Lab at Shanghai Jiao Tong University as a research assistant.
  • 2024 Started the M.Sc. program in Library and Information Studies at Hohai University.
  • 2024 Won the Excellent Work Award at the Intel Mini Hackathon for a fine-tuned LLM project.
Research

Research focus and background

I study trustworthy AI, building on my work in model evaluation and multimodal learning.

Trustworthy AI

My current focus is the trustworthiness of AI systems, with an interest in understanding their limitations and evaluating the reliability of their outputs.

Evaluation of code intelligence

My work on ClassEval-Pro and CodeOCR examines code generation and multimodal code understanding through benchmarks, test suites, and error analysis.

Multimodal credibility assessment

My work on review credibility combines textual, visual, and relational evidence to assess human-written and AI-generated reviews.

Publications

Selected publications

Selected publications and ongoing work.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding

Yuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen, Chenxu Zhang, Longfei Yun, Chengcheng Wan, Hongyu Zhang, David Lo, Xiaodong Gu

Proceedings of ISSTA 2026 · 2026

Studies code-as-image representations for multimodal code understanding and shows how visual encoding can improve efficiency while remaining competitive on downstream tasks.

Conference

ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation

Yeheng Chen*, Chaoxiang Xie*, Yuling Shi, Wenhao Zeng, Yongpan Wang, Hongyu Zhang, Xiaodong Gu

* Equal contribution / co-first authors (Yeheng Chen and Chaoxiang Xie)

Proceedings of AIware 2026, Benchmark & Dataset Track · 2026

Introduces ClassEval-Pro, a benchmark of 300 class-level code generation tasks across 11 domains, built through an automated three-stage pipeline with complexity enhancement, cross-domain class composition, and real-world GitHub code integration. Each task is validated by an LLM Judge Ensemble and test suites with over 90% line coverage. Experiments on five frontier LLMs under five generation strategies show that the best model reaches only 45.6% class-level Pass@1, while error analysis highlights logic and dependency errors as the main bottlenecks.

Conference

Multi-Detector Credibility Fusion Network: A Neural Architecture for Robust Multimodal Review Credibility Assessment

Chaoxiang Xie, Ming Li

International Journal of Intelligent Systems · 2026

Presents MDCFN, a multimodal architecture for robust review credibility assessment across textual, visual, and relational signals.

Under Review
Experience

Research and professional experience

My research experience and background in Python backend engineering.

Research

Oct. 2025 - Present Shanghai, China

Research Assistant

LLM for Software Engineering Lab (LLMSE), Shanghai Jiao Tong University

Advisor: Prof. Xiaodong Gu

  • Led the experimental design and implementation for CodeOCR, integrating LiveCodeBench and RepoQA with syntax-highlighted code rendering and visual context compression of up to 8x.
  • Built reproducible Python pipelines for GPT, Gemini, and Qwen-VL, including batch inference, rate-limit retries, structured results, statistical tests, and error analysis. Wrote the experimental section of the paper.
  • Evaluated three class-level code generation strategies for ClassEval-Pro. Designed the dependency, interface, logic, and integration error taxonomy and wrote the results analysis. Co-first author of the AIware 2026 paper.
Jun. 2023 - Present Nanjing, China

Independent Researcher

Institute of Management Science, Hohai University

  • Independently developed MDCFN for human-written and AI-generated fake review detection, covering data collection, model design, training, ablation studies, and manuscript preparation.
  • Fine-tuned a DeBERTa-v3 three-class model using 500,000 training and 20,000 validation samples, selecting weights by Macro-F1.
  • Built a multimodal dataset with 33k+ reviews and 50k+ images. Combined textual, temporal, visual, and relational features, achieving 98.77% F1 on the self-built dataset. The manuscript is under review at the International Journal of Intelligent Systems.

Education & Industry

Sep. 2024 - Present Nanjing, China

M.Sc. in Library and Information Studies

Hohai University

  • Research concentration: natural language processing. GPA: 87/100.
  • Relevant coursework: Business Intelligence Analysis and Mining, Advanced Information Retrieval, Machine Learning Applications.
Sep. 2018 - Jun. 2022 Nanjing, China

B.Sc. in Information Management and Information System

Hohai University

  • Graduated with GPA 85/100.
  • Relevant coursework: Statistics, Database Principles, Data Structures, Information Security, Calculus, Linear Algebra.
Aug. 2026 - Present

Agent Algorithm Intern

若界人工智能实验室

Jul. 2022 - Jun. 2024 Shenzhen, China

Software Engineer, Backend Infrastructure

Inspur Morning Cloud Technologies Co., Ltd.

  • Built shared Python backend frameworks for HCM Cloud using Tornado, Celery, Redis, and MySQL/openGauss. Delivered configurable data import/export, work alerts, and concurrency controls.
  • Designed metadata-driven import/export for business models, nested records, field mappings, attachments, and permissions, supporting batches of tens of thousands of records.
  • Implemented asynchronous jobs with Celery, WebSocket/Redis progress updates, detailed error reporting, and configurable business-key deduplication for safe retries.
  • Built scheduled and manual alert workflows with rule evaluation, recipient resolution, templates, and notification channels. Contributed to SSO, data-push conflict prevention, and a concurrency-related patent application.
  • Supported requirements from 100+ enterprises and production troubleshooting across SQL, Redis, Celery, and Linux/Kibana logs. Received the 2023 Inspur Group R&D Rising Star Award.
Projects

Selected projects

Selected work in language models and data analysis.

Experimental · TypeScript, Node.js, Pi, Ollama, MLX

Correction

  • Built a local prompt correction pipeline for coding agents. It edits natural language while preserving code, paths, commands, and numbers, with approve, edit, bypass, and cancel controls in the Pi terminal UI.
  • Implemented protected-span recovery, semantic checks, layered configuration, and JSONL diagnostics. Evaluated 200 frozen prompts and 1,000 protected-span samples across local models and 11 prompt versions.
  • In 300 time-limited validation runs, outcome p95 was 1.80 s; all 100 timeouts retained the original input. The project remains experimental, with long-input latency still being improved.

Released · Chrome Extension, Node.js, Readability, Turndown, lark-cli

Feishu Web Clip

  • Built a Chrome extension and local bridge to save web articles, images, and related PDFs to Feishu, with configurable AI summaries and streaming progress.
  • Implemented persistent clip attempts for timeout and partial-failure recovery, paired bridge credentials, an Origin allowlist, and minimal browser permissions. The extension reuses lark-cli authentication without storing Feishu access tokens.
  • Released on the Chrome Web Store and under active development.

Competition prototype · Private repository · Swift / SwiftUI, MNN-LLM, Node.js

FreeGo travel agent

  • Led a two-person team to build a travel planning and price-monitoring prototype for the Qwen on-device model application competition.
  • Designed on-device itinerary understanding and explanations, with cloud support for authorized data collection, scheduling, synchronization, and APNs notifications.
  • Implemented streaming model calls, tool-call parsing, search integration, quote validation, and a staged planning loop, with tests for the agent loop and synchronization boundaries.

Personal publishing workflow · Hermes Agent Skills, RSS, arXiv, GitHub Trending

Hermes news and WeChat publishing tools

  • Integrated 100+ RSS sources, arXiv, GitHub Trending, and URL summaries with concurrent fetching, caching, and persistent state.
  • Built a WeChat draft workflow with placeholder, frontmatter, credential, and dry-run checks. Supported daily publishing for 想开一间信心花舍 from July 27 through the August 17, 2026 reporting period.