Benchmarking & evaluation
Datasets, metrics, and realistic enterprise task environments for performance, safety, and reliability.
NeurIPS 2026 Workshop · AABA4ET
Where Agentic AI meets the real world of work
Benchmarking, evaluating, and deploying intelligent agents for complex enterprise operations at scale.
Accepted for NeurIPS 2026 — submissions open through August 31, 2026 (AoE).
About the workshop
Foster collaboration toward robust, efficient, and trustworthy Agentic AI for complex, dynamic enterprise operations.
The 2nd Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks continues the AAAI 2026 edition in Singapore. We connect cutting-edge agent research with the practical demands of evaluation and real-world deployment.
Focus areas
Original work across benchmarking, applications, safety, and multi-agent systems in enterprise settings.
Datasets, metrics, and realistic enterprise task environments for performance, safety, and reliability.
On-site understanding, planning, observation, reflection, and system management in business contexts.
Failure recovery, distribution shift, guardrails, and trust for long-running production agents.
Assistants that augment people inside real enterprise workflows.
Vision, text, and audio for robust decision-making in physical enterprise data.
Multi-agent and tool strategies for complex, multi-step enterprise goals.
Important dates
All deadlines are Anywhere on Earth (AoE).
Call for papers
Double-blind review. Dual submission welcome. At least one author must present in person.