Humanbound website
Evidence and Reporting

Turn every test into evidence you can hand to an auditor

Humanbound records every attack, verdict, and fix, then turns that record into branded reports, verified findings, and a posture score that show your board, your auditors, and your regulators how your AI agents were tested and how they held up.

Audit-ready reports

Four report levels, from an executive overview across every agent to the full transcript of each test conversation, all ready to print to PDF for formal submission.

Verified findings

Findings persist across test cycles, and a regression retest replays the original attacks against the current agent to confirm that a fix actually holds.

A posture score leadership can read

One score from 0 to 100 with an A to F grade, tracked over time and rolled up across every agent in your organisation.

The right depth for every reader

A board member needs to know whether your agents are safe to run, an engineer needs to know which conversation broke, and an auditor needs the conversations themselves. Humanbound generates branded HTML reports at four levels of scope from the same data.

Branded HTML reports

Four levels of scope, generated from the same underlying data.

Ready for auditors

A summary saying an agent was tested is a claim; the transcript of the attack and the reply is evidence. Every report includes methodology, a technology disclaimer and a legal notice, prints straight to PDF, and is built for submission under DORA, PCI-DSS, ISO/IEC 42001, NIS2 and the EU AI Act.

Organisation report

An executive overview across every project: organisation posture, findings by severity, and each agent's grade, score and monitoring status.

Project report

An agent's standing posture: permitted and restricted operations, findings by threat class, and 90 days of assessment history.

Assessment report

One test run in full: pass rate, posture before and after, and every conversation with its verdict and explanation.

Experiment report

One testing engine's run, with its methodology (OWASP or QA dimensions), metrics, vulnerabilities and conversations.

Findings that hold up under scrutiny

A finding marked fixed by hand only shows that someone believed it was fixed. Humanbound keeps findings as persistent records that carry their evidence, so ten conversations that expose the same weakness become one finding backed by stronger evidence.

Each finding counts the test cycles it has survived and moves through open, fixed, regressed and stale. Ownership is tracked from assignment to verification, with webhook events at each step, and findings export as JSON.

Loading...

A regression retest replays a finding's own recorded attacks against the current agent and reports one of three outcomes.

Not reproduced

None of the replayed attacks succeeded against the current agent, which is recorded as evidence that the fix holds.

Still vulnerable

At least one replayed attack succeeded, and the finding automatically moves from fixed back to regressed.

Insufficient evidence

The finding had no recorded attack to replay, so nothing could be tested, and the result is never treated as a pass.

Your obligations, tested

A rule against personalised financial advice is a boundary just like a rule against reaching admin functions. Humanbound attacks both with the same multi-turn engine.

Add your own policies to the restricted list. High-stakes domains get higher severity, and on the platform the restrictions apply to every cycle.

Loading...

One number your board can follow

Boards won't read transcripts, but they need to know whether your agents are getting safer or riskier. Humanbound sums up each agent's security in one score from 0 to 100 with a letter grade.

A

90 to 100
Excellent

B

75 to 89
Good

C

60 to 74
Fair

D

40 to 59
Poor

F

0 to 39
Critical

The score combines severity impact (active findings weighted by severity and status) with coverage effectiveness (the share of tested threat classes that pass), so a high score means few serious findings and broad coverage. Trends, a coverage breakdown and an organisation view show where it is heading.

Evidence from the first local run, and a full record on the platform

The open-source engine produces evidence from your first local run: hb report builds a branded HTML report and hb logs exports every conversation with its verdict. The platform keeps the record over time.

Open source

No login required
Run locally

Platform

The full record over time
Start Free

On every run

Posture score and grade
Posture score and grade
Branded HTML report for a test run
Branded HTML report for a test run
Conversation logs with verdict, severity, and explanation
Conversation logs with verdict, severity, and explanation
JSON export
JSON export

Over time, on the platform

Persistent findings with lifecycle and occurrence count
Persistent findings with lifecycle and occurrence count
Regression retests
Regression retests
Finding delegation and verification
Finding delegation and verification
Project, organisation, assessment, and experiment reports
Project, organisation, assessment, and experiment reports
Posture trends
Posture trends
Organisation posture
Organisation posture
Compliance restrictions applied to every monitoring cycle
Compliance restrictions applied to every monitoring cycle

Show how your AI agents were tested

Run your first test with the open-source engine and generate a report from it, or book a demo to see the full set of reports on the platform.