SOC 2 Type II audit in progress · EU, US and India data residency · a DPA you can sign

Trust Center

Methodology

How we measure

FlowTux runs code against your production systems, so a number on a marketing page is not a small thing. This page defines every metric we publish: what counts, what does not, where the figure came from, and what we do when we get one wrong. If a claim elsewhere on this site links here, this is the definition it is claiming against.

What “auto-resolved” counts

A ticket counts as auto-resolved only when all four of these are true. Anything short of this is counted as AI-assisted or human-handled, never as auto-resolved.

  • Tux AI set the category, priority, and owner without a person editing them.
  • The fix ran from the allow-list, or the answer was returned to the requester directly — no human took an action on the ticket.
  • The result was verified after the action ran, not merely dispatched.
  • The ticket closed and stayed closed — a reopen inside seven days removes it from the count.

Deduplicated tickets are counted once, against the surviving ticket. Collapsing fifty reports of one outage into a single ticket does not produce forty-nine auto-resolutions.

The allow-list

The device agent is default-deny. It can only run remediation actions from a fixed, allow-listed set of safe commands — flushing DNS or restarting a service, for example — and cannot execute arbitrary commands. Each action runs as a tracked job with an acknowledgement and a timeout, and every executed command is logged.

You approve that list. Most teams start narrow — password resets and access grants — and widen it as the audit log earns trust. Your auto-resolution rate is therefore a function of how much you have allow-listed: a team that permits three categories and a team that permits twenty will not see the same number, and neither figure is wrong. Full controls are documented on the security page.

What the AI never does unattended

Regardless of allow-list configuration, these always route to a human:

  • Any action against a system you have not explicitly allow-listed.
  • Anything requiring judgment about a person — access revocation, HR-adjacent requests, escalations naming an individual.
  • Novel failures with no documented fix, and any issue where verification of the fix is not possible.
  • Arbitrary or operator-supplied shell commands. The agent has no such capability to grant.

These tickets still get triaged, prioritised, deduplicated, and routed, and for code-adjacent issues the likely files are attached — so the human who picks one up starts with a diagnosis rather than a blank ticket.

How we measure

The auto-resolution figure on this site is 50% — the unweighted mean across every published customer. The individual results are 41–58%:

  • Stemlen 58% (B2B SaaS, 46 people, 3-person IT team)
  • SSK Grains Enterprises 41% (Agri-commodities trading, 210 employees across 5 locations, 2-person IT team)
  • Go Labs 52% (Developer tools, 28 people, no dedicated support team — engineers triaged directly)

We publish the mean rather than the best result, and we do not round it up. Your own rate depends on how many categories you allow-list, which is why the spread is this wide.

Every product figure we publish is pulled from the same workspace dashboards our customers see. We do not run a separate marketing analytics path, and we do not publish a number a customer could not reproduce from their own account.

  • Ranges, not averages of one. Where results differ by team, we publish the range and say which team profile sits where.
  • The company is always named.Every result on this site belongs to a company we name, with figures from their workspace. Quotes carry the speaker’s name wherever they have approved it; where they have approved the words but not their name, the quote shows their role alone and no name is invented to fill the gap.
  • Modelled figures are labelled at the source. Where we show a projection rather than a measurement, the chart says so and links here.

Third-party benchmarks

Some figures on this site come from published industry research — principally McKinsey (2026) and Freshworks (2025). These describe what AI-first support does across the category. They are not FlowTux measurements and we do not present them as ours; they appear only to establish that the category is moving, and the study is always named next to the number.

Corrections

If a number here is wrong, we change it and say what changed. If you think a figure on this site is overstated or a definition is doing convenient work, tell us at /contact and we will either show the derivation or correct the page.