We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon Office-suite tasks with economic grounding, spanning word-processing documents, spreadsheets, presentations, and cross-file productivity tasks. The benchmark is constructed from authentic Office requests proposed by practitioners and adapted through a privacy-preserving process that removes sensitive information while preserving user intent, constraints, and natural request phrasing. OmegaUse-OfficeVal further provides task-level economic grounding through two complementary signals. Human labor time records the time required by human workers to complete the task without LLM assistance, while task price proxy estimates the market price of completing the task using explicit price signals when available and expert estimates otherwise. These annotations enable evaluation in terms of both task completion and the human effort and economic value associated with the work. We construct the benchmark through a multi-stage task adaptation pipeline and evaluate final deliverables using deterministic code-based verifiers. Experiments with frontier LLM agents and general-purpose computer-use or coding agents show that current systems make only partial progress on these realistic Office-suite workflows. OmegaUse-OfficeVal offers a reproducible, economically grounded testbed for measuring progress toward reliable LLM agents for everyday Office productivity.
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on economically grounded, long-horizon office-suite tasks.
Task demos
Task Examples of OmegaUse-OfficeVal
Abstract
Long-horizon Office-suite tasks grounded in authentic economic demand.
Task overview
100 authentic tasks reflecting real long-horizon Office work.
What each task contains
Realistic inputs, human value, and executable evaluation.
Instruction
The natural-language request and required final deliverable.
Input files
Office documents, PDFs, images, video, or audio.
Human labor time
Quality-controlled human completion time without LLM assistance.
Task price proxy
An explicit market price or independent expert estimate.
Golden artifact
A professional ground-truth deliverable.
Code verifier
Executable rubric-based evaluation for the submitted artifact.
Evaluation focuses on the final artifact, allowing agents to use GUI actions, scripts, APIs, or hybrid workflows.
Task distribution
Real demand produces a deliberately unbalanced benchmark.
Construction pipeline
From 1,715 practitioner proposals to 100 verified Office tasks.
Practicality filtering and three-expert screening retain only authentic, feasible long-horizon work.
At least two raters measure each task; a third is added when completion times differ substantially.
Experts iteratively resolve human-code discrepancies until the checks align with task intent.
Leaderboard
Completion and economic-value weighted performance.
| Rank |
|---|
Scores are reported on the OmegaUse-OfficeVal benchmark. Use the controls to compare overall, time-weighted, and price-weighted performance.