AI Agents Can Pass Tests While Taking Risks

Miami AI startup Molt AI says conventional testing can miss harmful actions by AI agents. Its Fisher system tests what agents actually do, including tool calls, data retrievals and system changes.

August 13, 2026
AI Agents Can Pass Tests While Taking Risks Companies

Summary: Miami AI startup Molt AI is arguing that enterprises need to test what AI agents actually do, not simply what they say. Its Fisher assessment system analyzes tool calls, data retrievals and other actions, with the company reporting more than 50,000 adversarial test episodes and examples where agents appeared to refuse a request after already accessing sensitive data.

moltaiaiagentsaiassuranceaisecurityaiinfrastructureenterpriseaicybersecuritymiamistartups

An AI agent can give the right answer and still have done something wrong.

That is the problem Miami startup Molt AI says conventional AI testing is failing to catch.

Molt has launched Fisher, an agent-assurance assessment designed to test AI systems based not only on their final responses, but on what they actually do along the way: the tools they call, the information they retrieve, the APIs they access and the changes they make to connected systems.

The company has also raised $1 million in pre-seed funding and published a technical white paper documenting its testing methodology.

“A refusal can hide a completed action,” - Greg Frank, co-founder of Molt AI and creator of Fisher.

That distinction could become increasingly important as companies give AI agents access to databases, internal systems and other tools that allow them to act rather than simply generate text.

The Problem With Testing the Final Answer

Most people interacting with a chatbot see the conversation.

An enterprise AI agent can be doing considerably more behind the scenes.

An agent might query a database, retrieve a document, call an API or send a message before producing a final response to the user.

If an evaluation only looks at that final response, a system could appear to have handled a request correctly even if it had already crossed a security boundary.

Molt's argument is that agent testing therefore needs to examine the entire trajectory.

That includes tool calls, retrievals, writes, API activity and other observable actions.

Fisher Tests What Agents Actually Do

Fisher uses adaptive, multi-turn adversarial testing to probe enterprise AI agents.

Rather than relying solely on a predefined library of prompts, the system can adapt its testing strategy based on what happens during an interaction.

Molt says its findings are classified as exploratory, observed or confirmed. A potential weakness becomes a confirmed finding only when it can be reproduced under recorded replay conditions.

The company also says Fisher can test proposed fixes by repeating the original attack and, where authorized, launching a fresh adaptive attack.

That is intended to answer a more useful security question than simply whether a vulnerability was discovered:

Did the fix actually work?

More Than 50,000 Adversarial Episodes

According to Molt, Fisher had conducted more than 50,000 multi-turn adversarial episodes and more than 500,000 conversation turns across more than 20 model architectures as of July 2026.

The company reports that adaptive tactics and refusal pivots increased verified attack success by between 6.6 and 14.6 percentage points in matched experiments involving two model families and two scenarios.

Those are Molt's own results rather than an independent industry benchmark, but they illustrate the company's central argument: testing strategy can affect what security evaluations discover.

Molt also describes a controlled historical comparison involving a database agent in which a conventional evaluation marked several sessions as defended because the final response refused the request.

According to the company, the underlying action trace showed that password and API-key records had already been queried.

Molt explicitly limits that result to its particular target, test set and date rather than presenting it as a universal performance comparison.

Why Agentic AI Changes the Security Problem

The distinction between a chatbot and an AI agent is becoming increasingly important for enterprise security.

A chatbot can generate an incorrect answer.

An agent can potentially take an incorrect action.

That changes the consequences of failure.

As enterprises connect agents to internal databases, software systems, communication tools and business processes, security teams need visibility into what happens between the user's request and the agent's final response.

That is the market Molt is targeting.

“The moment an agent can act on a company's systems, testing it with adversarial prompt libraries or merely reading its final answer stops being enough,” said Walton Comer, Molt's co-founder and CEO.

What This Means For Miami

Molt is headquartered in Miami, with offices in New York, putting the company inside one of the city's expanding AI startup categories: enterprise infrastructure and security.

Miami's AI ecosystem has attracted companies working across fintech, healthcare, real estate and other industries where AI agents could eventually interact with sensitive business systems.

That creates a natural market for tools designed to test those systems before they are given significant autonomy.

Molt's $1 million pre-seed round will support expansion of Fisher, enterprise assessment and integration capabilities and further research into agent behavior and verified remediation.

The larger opportunity, however, goes beyond one startup.

As businesses move from AI that answers questions to AI that takes actions, the definition of a safe AI system may have to change with it.

A polite refusal at the end of a conversation is no longer necessarily evidence that everything went right.

The important question may be what happened before the agent said no.

Reporting Source: This article builds upon reporting from Molt AI's announcement and adds analysis of what the development means for Miami and South Florida.

← Back to MAIN