Benchmark — August 2026

Independent OWASP LLM Top 10 + content-moderation evaluation of 10 AI guardrail backends.

707 probes · 10 backends · generated 2026-08-04 13:05 UTC · run ac387219 · guardrailprobe v0.1.5

llm_guard
Winner — 85.9% overall
guardrails_ai
Best accuracy / latency ratio
−5.1%
Biggest regression — azure_content_safety
707
Total probes run · 0 backends skipped
Signed PDF JSON Markdown Methodology

Overall comparison

BackendOverallvs last monthBestWorstAvg latency
openai_moderation100.0%+0.0%LLM01LLM0133 422 ms
llm_guard85.9%+0.0%LLM01LLM102 569 ms
lakera83.3%+1.3%LLM01LLM10584 ms
nemo79.5%−5.1%LLM01LLM1031 454 ms
aws_bedrock59.0%+0.0%LLM01LLM10660 ms
azure_prompt_shields24.4%+0.0%LLM01LLM091 613 ms
azure_content_safety20.5%−5.1%LLM02LLM011 620 ms
llama_firewall11.5%+0.0%LLM01LLM058 043 ms
presidio6.4%+0.0%LLM02LLM01974 ms
guardrails_ai2.6%+0.0%LLM01LLM050 ms

Low scores are not always failures: Presidio is a PII-detection tool being tested against injection probes, and guardrails_ai's 0 ms latency means its hub validators never actually executed (see the methodology's two-tier categorisation of general-purpose vs specialised tools).

Per-category winners (OWASP LLM Top 10)

CategoryDescriptionWinnerScoreRunner-upScore
LLM01Prompt Injectionlakera100%aws_bedrock100%
LLM02Insecure Outputlakera81%llm_guard81%
LLM03Training Data Poisoninglakera100%aws_bedrock100%
LLM04Model DoSlakera67%llm_guard67%
LLM05Supply Chainllm_guard100%lakera80%
LLM06Sensitive Info Disclosurelakera100%llm_guard100%
LLM07Insecure Pluginlakera83%llm_guard83%
LLM08Excessive Agencylakera100%llm_guard100%
LLM09Overreliancellm_guard100%nemo100%
LLM10Model Theftlakera20%llm_guard20%

Content moderation

BackendHateViolenceSexualSelf-harmOverall CM
nemo100%100%100%100%100%
lakera100%100%100%100%100%
aws_bedrock100%100%100%100%100%
llm_guard100%100%100%100%100%
azure_content_safety80%100%60%80%80%
llama_firewall0%0%20%0%5%
guardrails_ai0%0%0%0%0%
presidio0%0%0%0%0%
openai_moderation0%0%0%0%0%
azure_prompt_shields0%0%0%0%0%

Accuracy vs latency

BackendOverallAvg latencyClassRecommended for
guardrails_ai2.6%0 msUltra-fastReal-time, high-throughput pipelines
lakera83.3%584 msModerateBatch processing, async pipelines
aws_bedrock59.0%660 msModerateBatch processing, async pipelines
presidio6.4%974 msModerateBatch processing, async pipelines
azure_prompt_shields24.4%1 613 msSlowOffline analysis, compliance audits
azure_content_safety20.5%1 620 msSlowOffline analysis, compliance audits
llm_guard85.9%2 569 msSlowOffline analysis, compliance audits
llama_firewall11.5%8 043 msSlowOffline analysis, compliance audits
nemo79.5%31 454 msSlowOffline analysis, compliance audits
openai_moderation100.0%33 422 msSlowOffline analysis, compliance audits

Month-over-month

BackendJulyAugustChangeStatus
lakera82.0%83.3%+1.3%stable
aws_bedrock59.0%59.0%+0.0%stable
azure_prompt_shields24.4%24.4%+0.0%stable
guardrails_ai2.6%2.6%+0.0%stable
llama_firewall11.5%11.5%+0.0%stable
llm_guard85.9%85.9%+0.0%stable
openai_moderation100.0%100.0%+0.0%stable
presidio6.4%6.4%+0.0%stable
azure_content_safety25.6%20.5%−5.1%regression
nemo84.6%79.5%−5.1%regression

Reproduce this benchmark

Any machine, ~10 minutes with credentials configured
pip install guardrailprobe
guardrailprobe run --year 2026 --month 8

Reports regenerate on the first of every month via GitHub Actions. Signed PDF carries an RFC 3161 timestamp — verify with guardrailprobe cert verify benchmark_2026_08.pdf.