Zhuohan Xie

About Me

I am a postdoctoral associate at Mohamed bin Zayed University of Artificial Intelligence (MBZUAI). My research focuses on trustworthy financial AI, especially benchmarks, shared tasks, and evaluation protocols for language and agent systems in high-stakes decision settings.

I use finance as a pressure test for domain intelligence: models must justify their reasoning, cite decisive evidence, respect financial and regulatory constraints, and remain stable when they act as agents. My recent work builds a connected line around FinChain, FinMMEval, SAHM, FinCARDS, and Herculean.

Research Interests

  • Verifiable financial reasoning: symbolic templates, executable checks, and domain constraints for auditable model reasoning.
  • Evidence-grounded analyst workflows: financial document QA, analyst reranking, reporting agents, and utility-grounded retrieval.
  • Financial agents: trading, hedging, auditing, memory stability, mandate drift, and risk-aware decision protocols.
  • Shared evaluation infrastructure: labs, leaderboards, final-test portals, working notes, and benchmark release discipline.

Collaboration

Open to focused collaborations that turn evaluation ideas into reusable public assets:

Verifiable finance benchmarks

Symbolic templates, executable checks, domain constraints, and metrics that expose reasoning failures.

Financial-agent protocols

Decision workflows, paper-trading windows, mandate drift, memory lock-in, and risk-aware diagnostics.

Shared-task infrastructure

Lab-style evaluation with participant systems, final-test portals, leaderboards, and working notes.

Analyst workflow benchmarks

Document QA, analyst reranking, reporting agents, recommendation, and grounded evaluation for finance.

News

  • FinChain and SAHM accepted to ACL 2026 Main as oral presentations.
  • FinCARDS, RealFin, and Same Claim, Different Judgment accepted to ACL 2026 Findings.
  • Conv-FinRe accepted to SIGIR 2026, extending the finance line into utility-grounded recommendation.
  • FinReporting accepted to ACL 2026 Demo.
  • Herculean released as an agentic benchmark for financial intelligence.
  • FinMMEval appeared at ECIR 2026 as evaluation infrastructure for AI in finance.

Selected Publications

Herculean workflow figure

arXiv 2026

Herculean: An Agentic Benchmark for Financial Intelligence

Xueqing Peng, Zhuohan Xie, Yupeng Cao, Haohang Li, Lingfei Qian, et al.

  • Benchmarks financial agents across trading, hedging, market insight, and auditing workflows.
  • Extends the finance line from static reasoning benchmarks to agentic decision protocols.

Portfolio

Financial AI work by program pillar

Each project is positioned as either a flagship anchor, a supporting asset, or an infrastructure component. The goal is cumulative ownership of a research line, not a loose pile of finance papers.

Verifiable financial reasoning

  • FinChainACL 2026 Main, Oral
  • SAHMACL 2026 Main, Oral
  • Culture-aware financial reasoningongoing

Evidence-grounded analyst workflows

Financial agents and decision protocols

  • Herculeanagentic financial intelligence
  • Agent mandate stabilityongoing
  • Strategic decision agentsongoing

Community evaluation infrastructure

  • FinMMEvalCLEF 2026 Lab / ECIR 2026
  • FinMMEval leaderboardsTask 1/2/3 results and live evaluation
  • Financial intelligence challenge planningcompetition and shared-task planning

Publications

Selected publications in AI for Finance

  1. FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning ACL 2026 Main, Oral · first author
  2. The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems ECIR 2026 · first author / organizer
  3. SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning ACL 2026 Main, Oral · last author
  4. Herculean: An Agentic Benchmark for Financial Intelligence arXiv 2026 · financial agent benchmark
  5. FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering ACL 2026 Findings · last author
  6. RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid? ACL 2026 Findings
  7. Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection ACL 2026 Findings
  8. FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures ACL 2026 Demo · last author
  9. Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation SIGIR 2026

Other Work

Adjacent trustworthy AI work

Earlier work in fact-checking, evidence retrieval, and text generation/evaluation supplies the broader trustworthy-AI foundation behind the finance program.

Experience

  • 2024.08 - Now: Postdoctoral Associate, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE.

Education

  • 2024.07: Ph.D., University of Melbourne.

Mentoring

I work closely with students and early-stage collaborators from problem framing to benchmark design, experiments, writing, rebuttal strategy, and release planning.

  • Rania Elbadry (July 2025 - present): SAHM, ACL 2026 Main Oral; The Geometry of Forgetting, arXiv 2026.
  • Yixi Zhou (January 2026 - present): FinCARDS, ACL 2026 Findings; SQLStructEval, arXiv 2026.
  • Shuzhi Gong (November 2025 - present): Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking, SIGIR 2026.
  • Fan Zhang (November 2025 - present): FinReporting, ACL 2026 Demo.

Services

I contribute to evaluation communities as a shared-task organizer and reviewer for NLP, machine learning, and trustworthy AI venues.

  • Organizer: CLEF 2026 FinMMEval Lab; SemEval 2025 Task 10.
  • Shared-task contributions: PAN / ELOQUENT, ImageCLEF, and financial AI evaluation infrastructure.
  • Reviewing: ACL Rolling Review, NeurIPS, COLM, EMNLP, ALTA, CHOMPS, TFAI, MisD, and SemEval 2027 task-proposal review committee.