top of page
agent check mockup breit4.png

Agent Check

agile product development  ·  agent evaluation tool  ·  2026

Agent Check is an AI-powered evaluation tool designed to help teams assess AI-agent ideas before a single line of code is written, predicting business value and surfacing risks along the way. It offers stakeholders a shared basis to align on the key questions and find out what's actually worth building.

Goal

Agent Check was the result of a 3 months programm where SAP set the challenge of how a pruduct could show the true value of an AI agent within a company before it is build.

Tools: Figma, Claude, TLDV, Miro + Frontend, Backend & DB Tools

Team: 4 (PM, Interaction Design, Development, AI Expert)

Time budget: 3 months

Everyone is Building AI Agents

But nobody can prove where they'll actually create value.

More than 80% of AI projects fail to deliver business value, not because the technology doesn't work, but because there's no shared way to define success before building starts. Developers ship fast, leaders approve budgets, and partners pitch ideas, while the gap between "this sounds great" and "this actually works" stays wide open.

Why do AI agent projects fail?
We conducted 14 problem interviews with SAP insiders, developers and consultants on how AI agents are integrated in companies
image 233.png
No Way to Prove Value

There's no standard framework to measure what an agent is worth — before or after deployment. Business decisions are made on gut feeling, costs are calculable but benefits aren't, and even when a score exists, nobody trusts it enough to act on it.

image 231.png
Build vs. Buy — and Whether to Use AI at All

Teams lack a clear framework for deciding when to use an agent, when simple automation is enough, and whether to build custom or use prebuilt solutions. Without that guidance, projects either overshoot or miss the mark entirely.

construction-sign_1f6a7.png
Pilots Don't Survive Production

Prototypes shine in demos, but scaling brings load, security, multi-user management, and regional permission complexity nobody planned for. Add legacy systems, inaccessible data, resistant employees, and the constant human supervision agents still need, and the promised efficiency quickly disappears.

image 232.png
Stakeholders Don't Speak the Same Language

IT, business, and HR all have different priorities, vocabularies, and risk tolerances. When an agent project spans all three, misalignment is almost guaranteed — and without a shared framework to communicate around, critical problems and risks get lost in translation.

The Solution? Agent Check
Get an evaluation on your AI agent idea before building to make sure it will bring true value to your company

1.

Explain your agent idea

2.

Get a custom evaluation

3.

Pitch it to your team

pretty agent check mockup.png
Challenges of Creating Agent Check
Frame 135.png
Measuring the Value of an AI Agent

A lot of research during this project went into finding out what a valuable AI agent would actually look like to different stakeholders. ROI is difficult to calculate upfront: token costs, usage rates, and actual time savings only become clear after deployment, if at all. Even efficiency gains are hard to predict: will employees actually adopt the tool, will it improve processes within the company or will supervising its output take just as long as doing the task by hand?

 

This is why Agent Check focuses on qualitative output rather than a single number. Instead of chasing a premature ROI figure, the tool surfaces the factors most likely to make or break a project. Will it scale? Will it be accepted? Is AI the right solution or can the problem be solved with automation?

 

The next step is making the tool more quantitative over time. By building a self-learning system that draws on accumulated evaluation data, Agent-Check can develop benchmarks across agent types, sharpen its predictions, and eventually help teams set realistic business goals with actual numbers behind them.

Balancing AI Autonomy and User Control

Finding the right balance between AI autonomy and user control was one of the core challenges of the project. Too little input led to AI slop; too much exhausted users and made them skip ahead.

 

Our solution minimized typing through speech input and a pre-analysis step, where the system reflected its understanding of the agent idea back to the user for a quick check and correction rather than starting from scratch. Making sure the user doesn't put effort into explaining something the tool could easily guess.

 

Also, verdicts were only shown where real user input existed. No hallucinations, no gap-filling. If a user wanted output on a specific topic, they were prompted to provide the missing input in exchange. This built trust through transparency: every output was traceable to what the user actually said.

user control image.png

Want to see more?

Would you rather...

bottom of page