
Summarize:
Testing is an interrogation. A reliable test does not merely ask whether the happy path works. It asks what the system does when evidence is noisy, inputs are hostile, dependencies fail, or the correct action is to refuse. The strongest AgentHack 2026 testing projects turn that idea into distinct, inspectable architectures. From deterministic probes and scoring around probabilistic components to explicit refusal paths, and failures captured as durable regression coverage.
This article focuses on what QA engineers can carry into their own builds. It covers four Test Cloud track winners plus SelfHeal QA, which won best cross-platform integration. The official announcement reports 1,400 builders, 333 submissions, 203 projects evaluated, and 31 finalists. This article follows the testing mechanics behind the results. Read the official winners announcement.
The projects occupy different points in one continuous quality loop. The diagram is intentionally a lifecycle view rather than a product architecture. It shows where each intervention creates the evidence that the next stage consumes.

QA takeaway: every probabilistic judgment needs a deterministic boundary around it. The boundary may be a reproducible fault, a score with named components, a browser observation, a defect record, or a regression test that can be rerun.
The judging bar matters because it explains why these are more than speculative patterns. AgentHack required a working solution running on UiPath Automation Cloud, a public GitHub repository with setup instructions and a license. A demo video of no more than five minutes showing the running solution and its architecture, and a presentation deck. External frameworks and LLMs were allowed, but UiPath had to be the orchestration and governance layer. See the published requirements.
For a testing audience, those requirements are useful acceptance criteria: can another engineer inspect the build, run the path, see where the agent is used, identify the human decision point, and trace a result into an auditable testing system? The projects below are valuable precisely where they make those boundaries visible and where they state their limits.
Jonathan SolvesProblems, AMER — $5,000 Architecture: a hybrid, flaky-test triage pipeline. FlakeWarden separates a deterministic scoring layer from an LLM classification layer. The deterministic component extracts repeatable signals from test history and behavior; the model then classifies the failure using that evidence rather than inventing a verdict from an unconstrained prompt. That separation is the architectural point: the model interprets bounded evidence, while the score remains inspectable and benchmarkable.
The tool is designed for the red-build decision: determine whether a failure is a genuine product regression or an unreliable test. On a labeled corpus of 150 failures, the submission reports 90.7% triage accuracy and a measured 0% safety-direction false-positive rate: no real regression was ever classified as flaky. About a third of those cases were settled by the deterministic scorer alone, with no model call at all; only the ambiguous middle band reached the Agent Builder classifier. A hard rule caps the score when a selector changes at the moment the test breaks, so a genuine UI regression cannot be waved off as flaky.
The external reference point is published research on flakiness at scale, not a shared benchmark dataset. Google has reported that roughly 16% of its tests showed some flakiness and that flaky failures sat behind about 84% of its red builds; those figures frame the problem in the submission rather than serving as its evaluation set. FlakeWarden’s accuracy is measured on the builder’s own labeled corpus, reproducible from the public repository. That is a real result, and a narrower claim than validation against an industry dataset.
For QA engineers, the pattern is portable. Keep evidence collection, feature calculation, thresholds, and release gates deterministic. Use the LLM where language-level synthesis helps—summarize a failure or assign a category—and preserve the raw signals, model output, confidence, and refusal path. A useful refusal is “insufficient history to classify”; it should route to human triage rather than silently turning uncertainty into “flaky.” View the FlakeWarden submission.
Kinnah Joshua, EMEA — $3,000
Architecture: a message-layer resilience test harness around a multi-agent workflow. A UiPath API Workflow loops through scenarios, calls a live Python service, and governs each verdict in Test Manager. The harness runs the target twice—once clean and once with a selected fault—then compares the outcomes and returns a 0–100 resilience score, verdict, and explanation.
The submission names six deterministic fault types: hallucinated output, refusal, dropped data, prompt injection, infrastructure crash, and infrastructure latency. In its reference healthcare prior-authorization pipeline, an injected, hallucinated approval turns a case that should escalate into a silent approval and scores 0, “Silent Corruption.” The same pipeline scores 70, “Mostly resilient,” against refusal and prompt injection because the resolution gate still escalates. The score is decomposed into outcome integrity, detection, and escalation, which makes it actionable rather than opaque.
The QA lesson is to test the handoffs, not only for each agent. Inject the same fault at the same boundary, compare it with a clean control run, and make “human escalation” a first-class assertion. Cricible also states an honest scope: one reference pipeline and one carefully chosen hard case, with live LLM reasoning that remains nondeterministic even though the injection is deterministic. View the Cricible submission.
Oleksandr Kryo, AMER — $2,000; also, Best First-Time Builder — $1,500
Architecture: an evidence-producing security validation agent. Penetron moves from scanner output to attempted exploit verification, narrowing a broad finding list into issues with demonstrated relevance. In testing terms, it turns “this might be vulnerable” into “this exploit path succeeded or failed under these conditions,” giving remediation a more useful basis.
The important boundary is evidentiary, not rhetorical. A failed exploit attempt does not prove a system is secure; it tells the team what was or was not demonstrated by the test. Store the request, target state, exploit trace, result, and limitations so a reviewer can distinguish coverage gaps from clean findings. View the Penetron submission.
Tim Dries, EMEA — $8,000
Architecture: A continuous adversarial evaluation system. FightArena is a Maestro Case that holds the long-running fight; RoundOrchestrator is the Maestro Flow for one round; a Python/LangGraph Red Coach powered by Claude Opus mutates attack strategies; the Blue target can be a UiPath Agent, Maestro Flow, or external LangGraph target; Test Manager stores winning attacks as persistent regression tests; Action Center holds human-approved remediation; and a UiPath Coded App provides the operator's surface.
The Red Coach is not just a generator of random bad prompts. It invents new attack personas when the existing corpus stops scoring hits, learns from successful defenses, and chains multi-turn fraud, prompt-injection, exfiltration, and tool-abuse strategies. The Devpost page reports 42 fights against the demo fixture “fake-ceo-naive,” with several attacks invented during the fights rather than seeded by the builder. Each breach is OWASP LLM Top 10, and MITRE ATLAS tagged and scored with the AVSS framework.
This is the clearest example of regression capture as an architectural output. The red team discovers; the judge scores; Test Manager preserves the breach; the fix recommender proposes a prompt patch, tool restriction, or guardrail; Action Center asks a human to approve the change; and the system can run the attack again. That feedback loop is the difference between a security demo and a maintainable quality practice.
It also has a meaningful implementation caveat: the submission lists the end-to-end RoundOrchestrator flow as a next step, while the demonstrated system includes the Maestro surfaces, Test Manager, Action Center, and deployed Coded App. Preserve that distinction in internal reporting. A compelling architecture still needs every claimed path to be demonstrably wired. View the Gauntlet submission.
Anthony Yanza — Best Cross-Platform Integration — $1,500
Architecture: a UiPath Coded Agent that selects tests based on changed files, drives a real browser with Playwright, triages a failure into a brittle locator versus a real bug, then either proposes a selector repair or refuses to heal and files a defect in Test Manager. The agent calls UiPath Identity and Test Manager APIs; Orchestrator runs the published package as a serverless job; Test Manager remains the system of record.
SelfHeal QA reports a genuinely brittle locator changing from #login-btn to #sign-in-btn and a benchmark with 0% false-negative rate across 8 real regressions plus 100% triage accuracy across a 16-case adversarial benchmark, including look-alikes. Those are the project’s reported benchmark results, not a general guarantee. The most important behavior is the refusal path: on a real bug, the agent does not patch around the failure; it creates a defect with reasoning attached.
For QA teams, this is the guardrail to copying. Define “heal” narrowly, prove the proposed change against the live page, and make “do not heal” a successful outcome when the evidence points to a product defect. A green test is not always a quality win; sometimes the highest-value result is a red test with a correctly filed defect. View the SelfHeal QA submission.
1. Separate deterministic control from generative judgment
Use deterministic fixtures, fault injection, feature extraction, score components, execution records, and release gates to make runs repeatable. Use generative models for exploration, synthesis, classification, or remediation of suggestions within those boundaries. Record both layers so a reviewer can reproduce the evidence and challenge the interpretation.
2. Treat refusal as a tested outcome
A safe system needs explicit “insufficient evidence,” “escalate,” “file a defect,” and “do not auto-approve” paths. Add them to the oracle and to the regression set. SelfHeal QA’s real-bug refusal, Cricible’s escalation, claim, and Gauntlet’s human approval step all make restraint observable.
3. Make every failure a durable artifact
A flaky classification should retain its signals and history. A resilience failure should retain the injected fault and clean comparison. A verified vulnerability should retain its exploit evidence. An adversarial breach should become a regression test. If the result disappears into a chat transcript, the organization has learned once and must rediscover it later. 4. Design for the tester’s review loop
4. Design for the tester’s review loop
The submission requirements offer a practical checklist: a runnable build, a public README, a short video that shows the system running and explains its architecture, and a deck that makes the orchestration and human role clear. Apply that standard internally. A test agent is production-shaped when another engineer can inspect its inputs, rerun its path, understand its verdict, and see what happens when confidence is low.
With UiPath Test Cloud, that future is becoming an enterprise-ready reality: agents and automations work across the testing lifecycle, while people remain responsible for the decisions that matter most—setting priorities, defining guardrails, and leading quality. This is the direction described in UiPath Dark Testing Factory white paper: testing becomes a self-operating quality system that plans, designs, executes, and maintains tests end-to-end, with human judgment moved to the places where it adds the most value. Explore agentic testing with UiPath Test Cloud .

Product Marketing Specialist, UiPath
Sign up today and we'll email you the newest articles every week.
Thank you for subscribing! Each week, we'll send the best automation blog posts straight to your inbox.
Sign up today and we'll email you the newest articles every week.
Thank you for subscribing! Each week, we'll send the best automation blog posts straight to your inbox.