Goal

Fuel iX Fortify is a B2B enterprise product for automated AI red-teaming — adversarial testing that surfaces security and safety vulnerabilities in AI models and agents. I led two connected research streams to support Fortify's product and go-to-market strategy: a UX competitive audit benchmarking how rival red-teaming tools present findings to users, and a market analysis mapping the competitive landscape against the newly published OWASP Agentic Top 10 framework.

Background

AI red-teaming is a fast-moving category, and the tools in it inherit a hard UX problem: the output is dense, technical vulnerability data that has to be understood by both security engineers and non-technical business stakeholders. At the same time, the risk landscape itself was shifting — the OWASP Agentic Top 10 had just been published to classify the distinct risks introduced by autonomous, tool-using AI agents, separate from single-model LLM risk. Fortify's leadership needed to know two things: how well competitors were solving the comprehension-and-actionability problem in their UI, and where the market's actual security coverage stood against this new framework — so Fortify could prioritize its roadmap and market position accordingly.

Methods

Participants

UX audit: 6 participants across 2 design-workshop sessions, reviewing the publicly available UIs of 6 direct AI red-teaming competitors.

Market analysis: desk research across 5 vendors spanning pure-play red-teaming tools and larger security platforms with red-teaming modules.


Study Design

The UX audit used a qualitative, workshop-driven approach. I ran two structured sessions where participants evaluated competitor products against two research questions — how effectively they visualize findings, and how well they help users act on them. Observations were deductively coded into five categories, then refined inductively as new themes emerged, and standout features were flagged separately. A 2×2 mapping exercise (organized/disorganized × easy/difficult to understand) let participants independently rank competitors and surface points of agreement and disagreement.

The market analysis mapped each competitor's publicly documented red-teaming capabilities against the 10 risk categories in the OWASP Agentic Top 10 (primary, secondary, or no coverage), cross-referenced against a third-party agent-risk scoring model, and translated into a business-facing "4 harms" framework so findings could be communicated to non-technical stakeholders, not just security teams.

Insights

The UX audit found that every competitor optimized for data completeness over comprehension — dense visualizations (some diagram types were called out as effectively unreadable), unexplained scoring logic, and recommendations that restated findings instead of guiding a fix. The clearest gap: none of the six tools closed the loop from "here's a vulnerability" to "here's what it means for your business and what to do about it" — a gap felt most acutely by non-technical stakeholders. One competitor stood out clearly for pairing a simple, guided interface with inline remediation actions, and became a useful UX benchmark for Fortify's own design roadmap.

The market analysis found that competitor coverage of the OWASP Agentic Top 10 clusters heavily around risk categories that resemble existing LLM and AppSec testing patterns — not necessarily the most dangerous categories. Several structurally novel agentic risks, like inter-agent communication and cascading multi-agent failures, were essentially untested industry-wide, largely because testing them requires a multi-agent harness nobody had built yet. That gap, combined with the finding that no competitor translates technical findings into business-impact language, pointed to concrete opportunities for Fortify to differentiate.

My Learnings

This project stretched me across two very different research modes in the same engagement — running hands-on qualitative workshops with real participants, and doing structured, framework-driven competitive research — and then synthesizing both into recommendations that two different audiences, design and product/marketing, could act on. I got much more comfortable translating dense technical security concepts into language a non-technical stakeholder or a business decision could use, which is a skill I think matters more as AI products get harder to explain to the people buying them.