MatToolBench Logo
MatToolBench
Materials Science Agent Benchmark
材料科学智能体基准测试

A comprehensive benchmark for evaluating LLM agents on real-world materials science software — spanning GUI tools, scientific code APIs, and data-retrieval interfaces.

一个用于评估大语言模型智能体在真实材料科学软件上表现的综合基准, 涵盖图形界面工具、科学代码接口和数据检索平台。

🧪 204 Tasks 🔬 10 Domains 🤖 3 Agent Types 📊 7 Models Evaluated 📄 MIT License
0+
Tasks
任务数
0
Domains
领域
0
Models Evaluated
评测模型
0
Tool Categories
工具类别
📌 What is MatToolBench?
📌 什么是 MatToolBench?

A Desktop Agent Benchmark for Materials Science

面向材料科学的桌面智能体基准

MatToolBench evaluates AI agents on real-world tasks across professional materials characterization and analysis software. It covers three agent types — GUI Agent (visual interaction), Origin Agent (script-based analysis), and Code Agent (programmatic database queries) — running inside a Windows 11 virtual machine environment powered by Docker & QEMU.

MatToolBench 在专业材料表征与分析软件上评估 AI 智能体的真实世界任务表现。 涵盖三种智能体类型:GUI Agent(视觉交互)、Origin Agent (脚本分析)和 Code Agent(代码数据库查询), 均运行于 Docker & QEMU 驱动的 Windows 11 虚拟机环境中。

🖥️

GUI Agent

图形界面智能体

Visually navigates materials software GUIs (Jade, Avantage, VESTA, DM, MatStudio) to complete characterization analysis tasks via screenshots and planned actions.

通过截图与动作规划,视觉操作材料科学软件 GUI(Jade、Avantage、VESTA、DM、MatStudio)完成分析任务。

~100 tasks
📊

Origin Agent

Origin 脚本智能体

Writes OriginPro scripts to process experimental data (XRD, XPS, Raman, Cycle, Step, CE) and generate publication-quality plots automatically.

编写 OriginPro 脚本处理实验数据(XRD、XPS、拉曼、循环曲线等)并自动生成图表。

~16 tasks
💻

Code Agent

代码数据库智能体

Generates Python code to query materials databases (MP, OQMD, PyMatgen, OPTIMADE) and retrieve structural and thermodynamic properties.

生成 Python 代码查询材料数据库(MP、OQMD、PyMatgen、OPTIMADE),检索结构与热力学性质。

~70 tasks
▶ Demo
▶ 演示视频

Watch MatToolBench in Action

观看 MatToolBench 演示

See how LLM agents navigate professional materials science software, write analysis scripts, and query scientific databases to complete complex real-world tasks. ⚡ Sped up

观看 LLM 智能体如何操作专业材料科学软件、编写分析脚本, 并查询科学数据库来完成复杂的真实世界任务。 ⚡ 视频已加速

📊 Task Statistics
📊 任务统计

Complete Task Statistics Across All Domains

全领域任务统计概览

MatToolBench comprises 204 tasks across 11 domains and 3 agent types. GUI tasks span 5 professional software tools and are graded by sub-criteria count — Easy (1–3 pts), Medium (4 pts), Hard (≥5 pts). Origin tasks cover 6 plot types (XRD, XPS, Raman, FTIR, Battery Cycling, Free Energy) with no difficulty split. Code tasks evaluate multi-property database queries across MP, OQMD, OPTIMADE, and Pymatgen.

MatToolBench 共包含 204 个任务,覆盖 11 个领域和 3 种智能体类型。 GUI 任务跨越 5 款专业软件工具,按子标准数量分级:简单(1–3 分)、中等(4 分)、困难(≥5 分)。 Origin 任务涵盖 6 种图表类型(XRD、XPS、拉曼、FTIR、电化学循环、自由能),无难度区分。 Code 任务评测 MP、OQMD、OPTIMADE 和 Pymatgen 的多属性数据库查询。

Table 1. Statistics of all benchmark tasks across categories and domains.
表 1. 所有基准任务的类别与领域统计。 E / M / H(仅 GUI):简单 / 中等 / 困难, 依据子标准数量划分(1–3 / 4 / ≥5 分)。 Code 任务的"均分"表示每任务的键值/文件匹配标准数。
Category类别 Domain领域 #Tasks任务数 E / M / H Avg Pts均分 Total Pts总分
GUI图形界面 Avantage 2014 / 2 / 4 3.2565
DM 2013 / 5 / 2 3.3567
JADE 2015 / 5 / 0 2.9559
MS 206 / 11 / 3 4.1082
VESTA 2013 / 2 / 5 3.4569
Origin绘图 Origin 16— 2.0032
Code代码查询 MP OQMD OPTIMADE Pymatgen 80— 180
Mixed混合 Mixed 8— — —
Easy (1–3 pts, GUI)简单(1–3 分,GUI) Medium (4 pts, GUI)中等(4 分,GUI) Hard (≥5 pts, GUI)困难(≥5 分,GUI)
MatToolBench CLI evaluation pipeline
Figure 1图 1 : CLI output of the automated evaluation pipeline running agent tasks. Click to enlarge. :自动化评估流水线运行智能体任务的 CLI 输出。点击放大。
🔍 Benchmark Tasks
🔍 基准任务

Real-World Task Examples & Distribution

真实世界任务示例与分布

XRD Task Example
Figure 2图 2 : Example GUI task — the agent loads XRD data, identifies diffraction peaks, matches phases from a database, and plots the overlay pattern across 6 sequential steps. Click to enlarge. :GUI 任务示例——智能体依次加载 XRD 数据、识别衍射峰、从数据库匹配物相,并绘制叠加图谱,共 6 个连续步骤。点击放大。
Task Category Distribution
Figure 3图 3 : Task category distribution across all 204 benchmark tasks. Click to enlarge. :204 基准任务的类别分布。点击放大。

Task Category Breakdown

任务类别细分

Tasks span a wide range of materials science workflows, from crystal structure visualization and XRD/XPS spectral analysis to database queries for thermodynamic properties. Each task is graded by domain-specific sub-criteria, enabling fine-grained evaluation beyond simple pass/fail.

任务覆盖材料科学的多种工作流程,从晶体结构可视化、XRD/XPS 谱图分析, 到热力学性质的数据库查询。每个任务按领域特定子标准进行评分, 支持超越简单通过/失败的细粒度评估。

  • Structure Analysis (Avantage, VESTA, Jade)结构分析(Avantage、VESTA、Jade)
  • Microscopy (Digital Micrograph, MatStudio)显微镜(Digital Micrograph、MatStudio)
  • Data Processing (Origin XRD / XPS / Raman)数据处理(Origin XRD / XPS / 拉曼)
  • Database Queries (MP, OQMD, PyMatgen, OPTIMADE)数据库查询(MP、OQMD、PyMatgen、OPTIMADE)
🎬 Reference Trajectories
🎬 参考操作轨迹

Step-by-Step Human Reference Trajectories

逐步 人类参考操作 回放

Browse human-recorded reference trajectories for selected GUI tasks across five materials science tools. Each step captures the correct desktop action sequence for completing the instruction, serving as the ground-truth basis for automated evaluation accuracy.

浏览五款材料科学软件上选定 GUI 任务的人类示范操作轨迹。 每一步均记录完成该指令的正确桌面操作序列,作为自动化评测准确率校验的标准参照。

TASK
Trajectory screenshot
Step 1 / 1

The complete reference trajectory dataset — 100 human-recorded GUI instruction trajectories across five materials science tools — is available on Google Drive.

完整参考轨迹数据集(五款材料科学软件共 100 条人类 GUI 操作指令轨迹)已托管于 Google Drive。

Download on Google Drive 在 Google Drive 下载
🏗️ Benchmark Design
🏗️ 基准系统设计

System Overview & Evaluation Design

系统概览与评测设计

Tasks are loaded from JSON config files and dispatched to one of three agent pipelines. Each agent interacts with a Windows 11 VM (via QEMU inside Docker) or an external database API, and results are evaluated by domain-specific getters and metrics.

任务从 JSON 配置文件加载,分发到三种智能体流水线之一。 每个智能体通过 QEMU(Docker 内)与 Windows 11 虚拟机或外部数据库 API 交互, 结果由特定领域的 getter 和评分指标进行评估。

MatToolBench System Architecture
Figure 4: MatToolBench system architecture — task dispatch, agent pipelines, and evaluation flow. 图 4:MatToolBench 系统架构——任务分发、智能体流水线与评估流程。
📐 Evaluation Design
📐 评测设计

Each domain uses a tailored combination of metric functions. GUI domains rely on exact_match with specialized getter functions; Code domains use partial-credit matching (detect_kv_match) or binary file comparison (detect_file_match).

每个领域使用定制化的评测指标组合。GUI 领域使用 exact_match 配合专用 getter 函数; Code 领域使用部分分匹配(detect_kv_match)或二进制文件比对(detect_file_match)。

Table 2. 表 2. Metric functions per domain. ✓ = all tasks; fractions = partial usage. 各领域评测函数。✓ = 全部任务;分数 = 部分使用。
Domain领域 exact_match detect_kv detect_file
GUI + Origin
Avantage · DM · JADE
MS · VESTA · Origin
✓——
OPTIMADE✓——
MP—14/206/20
OQMD—✓—
Pymatgen——✓
Table 3. 表 3. Getter function types per GUI/Origin domain. GUI / Origin 领域的 getter 函数类型数。
Domain领域 #Getter Types种类数
Avantage6
DM6
JADE8
MS5
VESTA8
Origin2
🏆 Results
🏆 评测结果

Model Leaderboard

模型榜单

Accuracy Results

准确性结果

Score: normalized task score averaged over sub-criteria. SR (%): success rate (all sub-criteria satisfied).

Score:子标准归一化平均得分。 SR (%):成功率(所有子标准均满足)。

99 Best per domain领域最优
99 2nd per domain领域次优
99 Best overall总体最优
99 2nd overall总体次优
— Not evaluated待评测
Model 模型 Overall 总平均 GUI Agent Origin Agent Code Agent
Avg 均值 AvantageJADEDMMSVESTAAvg Origin PymatgenMPOQMDOPTIMADEAvg

Efficiency Results

效率结果

Steps: mean steps across all episodes (max 50, lower = better). Steps*: mean steps on successful episodes only. FTR (%): false termination rate (lower = better).

步数:所有回合平均步数(满分50步,越少越高效)。 步数*:仅成功回合平均步数。 FTR:错误终止率(越低越好)。

99 Best per domain领域最优
99 2nd per domain领域次优
N/A No successful episodes (Steps*) / no early terminations (FTR)无成功回合(步数*)/ 无提前终止(FTR)
Model 模型 Overall 总平均 GUI Agent Origin Agent Code Agent
Avg 均值 AvantageJADEDMMSVESTAAvg Origin PymatgenMPOQMDOPTIMADEAvg
🔬 Ablation Study
🔬 消融实验

Ablation Experiments

消融实验分析

Task Type Parameter Value Condition Template Script Workflow Hint Env Setup† Injected Content
GUI gui_hint_mode hint full — ✓ — GUI workflow instructions
no_hint baseline — ✗ — generic agent guidelines only
Code code_hint_mode hint full — ✓ — API usage examples & docs
no_hint baseline — ✗ — generic agent guidelines only
Origin origin_mode ×
origin_hint_mode
script + hint full ✓ ✓ ✓ domain template + OriginLab workflow
no_script + hint partial ✗ ✓ ✓ OriginLab workflow only
no_script + no_hint baseline ✗ ✗ ✗ generic agent guidelines only

† Env Setup: whether the task environment pre-opens Code Builder (Alt+4) before the episode begins. In no_hint mode this step is suppressed so the agent must independently discover the workflow tool, providing a clean baseline.

任务类型 参数 取值 条件 模板脚本 工作流提示 环境初始化† 注入内容
GUI gui_hint_mode hint 完整 — ✓ — GUI 工作流指令
no_hint 基线 — ✗ — 仅通用 Agent 指令
Code code_hint_mode hint 完整 — ✓ — API 使用示例与文档
no_hint 基线 — ✗ — 仅通用 Agent 指令
Origin origin_mode ×
origin_hint_mode
script + hint 完整 ✓ ✓ ✓ 领域模板脚本 + OriginLab 工作流
no_script + hint 部分 ✗ ✓ ✓ 仅 OriginLab 工作流提示
no_script + no_hint 基线 ✗ ✗ ✗ 仅通用 Agent 指令

† 环境初始化:任务开始前是否通过 setup 脚本自动打开 Code Builder(Alt+4)。在 no_hint 模式下该步骤被跳过,Agent 需从 Origin 主界面独立发现操作工具,从而构成干净的基线对照。

🖥 GUI & Code Tasks
🖥 GUI 与 Code 任务
Seed 1.8 Success Rate
GPT-5.4 Success Rate
hint (Seed 1.8) hint(Seed 1.8)
no_hint (Seed 1.8) no_hint(Seed 1.8)
hint (GPT-5.4) hint(GPT-5.4)
no_hint (GPT-5.4) no_hint(GPT-5.4)
performance gap 性能差距

Origin Tasks Origin 任务
📊 Origin Tasks — Three-Level Ablation
📊 Origin 任务 — 三级消融
Claude Sonnet 4.6 SR Score + Aesthetic Score  (max 32 pts) SR得分 + 审美分数  (满分 32 分)
SR Score — file successfully drawn (max 16 pts) SR得分 — 图像成功生成(最高16分)
Aesthetic Score — plot quality (max 16 pts) 审美分数 — 绘图质量(最高16分)
📐 Evaluation Quality
📐 评测质量

Evaluator Accuracy

评测脚本准确率

Automated evaluation scripts independently verify every numerical extraction and software-operation judgment. We report Accuracy, Precision, Recall, and F1 for both the Score (numerical value correctness) and Software (operation identification) dimensions.

自动化评测脚本对每项数值提取及软件操作识别进行独立验证, 分别从 数值得分(Score)和软件操作(Software)两个维度报告 准确率、精确率、召回率及 F1。

Colour key: 颜色说明: 1.00 — perfect 1.00 — 满分 0.95–0.99 — near-perfect 0.95–0.99 — 接近满分 ≤ 0.94 — notable gap ≤ 0.94 — 明显差距

Overall Performance

整体表现

Macro-averaged across all 20 tasks per software.

每款软件 20 个任务的宏平均值。

Software Score Software
— Acc.Prec.Rec.F1 Acc.Prec.Rec.F1
JADE 1.00 1.00 0.99 0.99 1.00 1.00 1.00 1.00
MS 1.00 0.99 1.00 0.99 1.00 1.00 1.00 1.00
Avantage 0.99 0.98 0.98 0.98 1.00 0.98 1.00 0.99
VESTA 0.98 0.94 0.99 0.97 0.99 0.94 1.00 0.97
📈 Origin Gallery
📈 Origin 图谱画廊

Representative Origin Plot Showcase

代表性 Origin 图谱展示

A curated stream of representative materials-science plots processed in Origin, spanning diffraction, spectroscopy, electrochemistry, thermal analysis, and energy-profile visualization.

精选展示使用 Origin 处理的代表性材料科学图谱,涵盖衍射、光谱、电化学、热分析与能量示意等多种类型。

XRD Pattern
XRD Pattern
XPS Spectra
XPS Spectra
FTIR Absorbance
FTIR Absorbance
Raman Spectrum
Raman Spectrum
Battery Cycling
Battery Cycling
XPS Fitted Spectra
XPS Fitted Spectra
FTIR Full Spectra
FTIR Full Spectra
Raman Mapping
Raman Mapping
Voltage-Time Curve
Voltage-Time Curve
FTIR Zoomed Spectra
FTIR Zoomed Spectra
Free Energy Diagram
Free Energy Diagram
Raman Grayscale
Raman Grayscale
📊 Origin Evaluation
📊 Origin 质量评估

Origin Plotting Quality Scores

Origin 绘图质量评分

Five evaluators — human users and four frontier LLMs — rate outputs from Ours and four baselines on a 1–5 scale across four spectroscopy / diffraction task types. Ours axis is highlighted in blue; larger area = better quality.

人类用户和四个前沿大模型对 Ours 及四种基线方法进行 1–5 分评分,涵盖四类光谱/衍射任务。 Ours 轴以蓝色高亮显示,面积越大代表质量越高。

Raman
XRD
XPS
FTIR