OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on economically grounded, long-horizon office-suite tasks.

Task demos

Task Examples of OmegaUse-OfficeVal

Six OmegaUse-OfficeVal task examples showing instructions, input files, task price proxy, human labor time, and golden Office artifacts.
Representative tasks preserve concise practitioner instructions while pairing them with realistic inputs, economic signals, and professionally completed golden artifacts. View full-size figure

Abstract

Long-horizon Office-suite tasks grounded in authentic economic demand.

We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon Office-suite tasks with economic grounding, spanning word-processing documents, spreadsheets, presentations, and cross-file productivity tasks. The benchmark is constructed from authentic Office requests proposed by practitioners and adapted through a privacy-preserving process that removes sensitive information while preserving user intent, constraints, and natural request phrasing. OmegaUse-OfficeVal further provides task-level economic grounding through two complementary signals. Human labor time records the time required by human workers to complete the task without LLM assistance, while task price proxy estimates the market price of completing the task using explicit price signals when available and expert estimates otherwise. These annotations enable evaluation in terms of both task completion and the human effort and economic value associated with the work. We construct the benchmark through a multi-stage task adaptation pipeline and evaluate final deliverables using deterministic code-based verifiers. Experiments with frontier LLM agents and general-purpose computer-use or coding agents show that current systems make only partial progress on these realistic Office-suite workflows. OmegaUse-OfficeVal offers a reproducible, economically grounded testbed for measuring progress toward reliable LLM agents for everyday Office productivity.

Task overview

100 authentic tasks reflecting real long-horizon Office work.

100 Office tasks
37 Word
39 PowerPoint
19 Excel
4 PDF
1 Word + Excel

What each task contains

Realistic inputs, human value, and executable evaluation.

01

Instruction

The natural-language request and required final deliverable.

02

Input files

Office documents, PDFs, images, video, or audio.

03

Human labor time

Quality-controlled human completion time without LLM assistance.

04

Task price proxy

An explicit market price or independent expert estimate.

05

Golden artifact

A professional ground-truth deliverable.

06

Code verifier

Executable rubric-based evaluation for the submitted artifact.

Evaluation focuses on the final artifact, allowing agents to use GUI actions, scripts, APIs, or hybrid workflows.

Task distribution

Real demand produces a deliberately unbalanced benchmark.

Sunburst chart of 100 OmegaUse-OfficeVal tasks by professional domain and operation type.
Inner ring: professional domain. Outer ring: requested operation. View full-size figure

Construction pipeline

From 1,715 practitioner proposals to 100 verified Office tasks.

OmegaUse-OfficeVal construction pipeline: practitioner task collection and screening, economically grounded value estimation, and iterative code-based verification.
The benchmark combines authentic task collection, privacy-preserving artifact reconstruction, human labor and price annotation, and expert-aligned executable verification. View full-size figure
Task funnel 1,715 → 595 → 282 → 100

Practicality filtering and three-expert screening retain only authentic, feasible long-horizon work.

Economic grounding 20 qualified annotators

At least two raters measure each task; a third is added when completion times differ substantially.

Reproducible scoring Rubrics + code verifiers

Experts iteratively resolve human-code discrepancies until the checks align with task intent.

Leaderboard

Completion and economic-value weighted performance.

Primary metric
OmegaUse-OfficeVal leaderboard
Rank

Scores are reported on the OmegaUse-OfficeVal benchmark. Use the controls to compare overall, time-weighted, and price-weighted performance.