Humanbound website
Runtime Defense

Judge every input before it reaches the model.

The firewall sits on every path into your agent and decides what may direct it, what may only inform it, and what has been tampered with.

Every path classified

User messages, tool outputs, documents, and the agent's own records are each judged by who wrote them and what they're allowed to do.

Most requests never reach the LLM

Four tiers trade speed for depth. Clearly legitimate and clearly malicious inputs are settled locally at zero LLM cost.

Learns from your own tests

A classifier trained on your agent's adversarial test data catches attacks that generic detectors miss, and improves with every test cycle.

Judged by who wrote it, not where it came from.

Everything that reaches the model has the same authority. The model reads all of it as text and follows whatever looks like an instruction.

The firewall judges each payload by its author, not by the channel it arrived through. One policy file describes your agent. It never has to describe an attack.

Requests may direct the agent within policy. Ingested content may inform it but not direct it. Recalled records should only ever be data, and one that gives instructions has been tampered with.

Four tiers. Most requests never reach the LLM.

A prompt injection written as a customer support ticket will pass any regex or keyword filter. Catching it takes a language model, but running one on every input is slow and expensive.

Input enters at Tier 0 and moves up only when a lower tier can't reach a confident decision. Most requests are settled locally at zero LLM cost. Only the ambiguous ones reach the judge. Each verdict is pass, block with a named reason, or review, and you choose whether uncertain verdicts fail open or closed.

Every earlier turn counts as context, tool outputs included, so a multi-step attack is caught as a pivot.

Tier 0: Sanitization

No model call and zero cost. Blocks input that carries invisible control characters, zero-width joiners, or bidirectional overrides. Always active, and nothing is rewritten.

Tier 1: Attack detection ensemble

Pre-trained detectors such as DeBERTa, Azure Content Safety, or your own APIs run in parallel with a consensus threshold you set. Runs locally at zero cost and catches most known prompt injection patterns.

Tier 2: Agent-specific classification

A classifier fine-tuned on adversarial test data from your own agent. It catches attacks that generic models miss and fast-tracks requests that match known benign patterns. Also local and zero cost.

Tier 3: LLM judge

Full contextual analysis against your agent's security policy, permitted intents, and restricted actions. Called only when the lower tiers can't decide. Works with OpenAI, Azure OpenAI, Claude, and Gemini.

Every test cycle makes it sharper.

Generic detectors catch generic attacks. The attacks that get through are the ones shaped to your agent's specific behavior, its scope, its tone, the things it tends to comply with. Tier 2 closes that gap by learning from your own test logs.

1

Failed adversarial conversations, where the agent was compromised, become attack examples.

Passed QA conversations become benign examples. Run hb firewall train and it curates the data, trains the classifier, and saves a portable .hbfw model file.

2

Retrain after each test cycle.

Each run adds attacks and benign turns the previous model never saw, which means more coverage, fewer Tier 3 calls, and lower cost. You can also merge in results from PyRIT and PromptFoo.

What a real attack looks like. Two minutes, end to end.

PriceWatch, a repricing agent, is asked to check one price against a competitor. The competitor's page carries a bait link. The attacker's page behind it tells the agent to encode the company's cost, floor, and margin and send them off before answering. The agent complies, receives a fake price, and hands back a clean, confident recommendation.

Every byte of that attack entered through the ingest boundary. The user asked a normal question, so a filter that only checks user input never had anything to catch.

The demo agent is open source: github.com/dgerog/pricewatch-injection

Two lines on LangChain. One call anywhere else.

On LangChain, firewall.adapt_to("langchain") attaches the firewall to every boundary. For any other agent, call inspect() wherever content enters the model's context.

Pick which trust classes to enforce, and set block, log, or passthrough mode per class. The firewall itself uploads nothing.

Loading...

Put a firewall on every path into your agent.

Open-source Python library, Apache-2.0 licensed. Free to use, modify, and embed in commercial products.