Rules, checks and verification flows — contract
This is the implementation contract of abusend's decision logic: how a
verdict is decided, how checks (verification tools) are run and how the SDKs
drive them. Backend, SDKs, panel and docs are built against this file; change
it first when the contract changes. The customer-facing guide is
INTEGRATION.md; the edge API schema is api/openapi.yaml
(served at /v1/openapi.yaml); the product story is docs/PIVOT.md.
1. Concepts
| Concept | Meaning |
|---|---|
| Action | A protected flow (login, topup): a label with token required · optional and token_missing review · deny. Discovered from traffic (or added by hand). Actions have no mode: everything below applies on every action, configured or not. |
| Signal | Something abusend measures: TLS/JA4/H2, probe, browser, IP (datacenter, VPN, Tor, relay, country, ASN), device, velocity. Signals are rule fields and hints, not verdicts. |
| Risk | A derived signal: risk.score 0–100 and risk.level low · medium · high · critical · unknown from the built-in weighted signal definitions (internal/engine/signals.go, weights editable per project) and the project's level thresholds (§2.4). It never decides by itself. |
| Context | What the customer's server sends: user.attributes (persistent) and context (this request: amount, currency, plan). Typed. |
| Rule | when <condition> → allow · review · deny · challenge(check group); ordered; optional action scope; state Off · Monitoring · Enforcing (enabled + mode monitor · enforce). A rule's state is the only monitor/enforce switch: an enforcing rule decides in every action of its scope, including one seen for the first time; a monitoring rule only reports what it would do. |
| Check | A connected verification tool. Widget: turnstile, recaptcha (v3 / v2 invisible / v2 checkbox), recaptcha_enterprise, hcaptcha, friendly_captcha, geetest — run by the browser SDK (invisible, or visible in the SDK's verification dialog), verified by the edge — and abusend's own abusend_pow · abusend_hold · abusend_trace (no site key, no secret, no third party: the edge verifies them itself, §3.1). App: sms, email, kyc, custom — run by the customer's app, reported by its server. |
| Check group | What a challenge rule asks for: 1–3 distinct checks and a mode — any (one of them, picked at random by weight when it is issued), sequence (all, one after another) or parallel (all at once). A single check is a group of one (any). See §2.3. |
| Challenge | One issued verification of a check: pending → passed · failed · incomplete · error (a pending challenge past its expiry is incomplete). |
| Flow | Verdict → challenge(s) → final verdict of one attempt. At most 3 challenges, counting every issued one (a parallel group's included); each check once. |
Vocabulary (API → UI): allow Allow · review Your app decides ·
deny Block · challenge + check Verify with ‹check› (groups:
"Verify with Captcha or hCaptcha (70/30)", "… with Captcha, then SMS", "…
with Captcha and SMS at once") · check
Check · challenge Verification · action Action · rule states
Monitoring / Enforcing / Off (rules only) · signal Signal (hint).
Developer-facing names use challenge; "Verification" is only a UI label.
2. Verdict evaluation
2.1 Inputs
The token's assessment (signals), the user (stored attributes merged with the request's), the context, the presented challenges (if any) and their flow, the project's rules in priority order (with their states), the action's token requirement and the verdict time.
2.2 Built-ins, before any rule
-
Token:
reused,user_mismatch,invalid,project_mismatch,expired→ deny (decided_by.type = token), whatever the rules' states. The redemption is one atomic script on the token's hot record in Valkey (DESIGN §5.2): the first user to redeem binds the token, a repeated verdict is a retry only within 5 s of the first redemption, at most 3 redemptions, for the same action andclient_ip;user_mismatchwins overreused. A token whose record is gone (its token expired, or Valkey lost it) isinvalid; a state store that cannot be reached fails the verdict (503 state_unavailable), it never skips the check. Durability, a documented weakening of single use: Valkey persists with an append-only file synced every second (appendfsync everysec), so a redemption made in the last second before a Valkey crash can be lost and the same token redeemed once more within its lifetime (at most its 10 minutes by default, 1 hour at most); a record that was lost isinvalidinstead (the safe side). A token is still never valid twice at the same moment and a lost record never turns an unusable token valid. -
Block lists (users, IPs, devices) → deny, whatever the rules' states.
block_devicesapplies to the device id the browser presented (or was issued), never to a device recognition merely linked it to (§2.6). -
Token missing: on a
token: requiredaction withtoken_missing = deny→ deny. Withreview(the default) the rules still run (a withheld token never skips a check or a deny rule) and the token only floors their outcome: where they would allow, the decision is review decided bytoken(the would-decision is floored the same way). On atoken: optionalaction the rules run withtoken.status = "missing",probe.status = "none",risk.level = "unknown"(IP signals come fromclient_ip). -
Presented
challenge_ids(0–3ch_…ids; the SDKs forward the comma-separatedX-Abusend-Challenge; more than 3 or a malformed id → 400invalid_request). Each id is checked on its own:- unknown, other action or user → no credit (info reason
challenge_invalid); - issued with a device (a valid token) but the retry has no valid token or another device → no credit;
- consumed by the open-challenge lookup (another tab) → no credit;
- consumed by a verdict: within 5 s by the same caller (user,
client_ip, device, action) → the stored response is replayed; otherwise denychallenge_reused; - passed more than 5 minutes ago and unconsumed → no credit (use-by).
The remaining ids must belong to one flow: the flow of the latest-created presented challenge; ids of other flows give no credit (
challenge_invalid). That flow is attached, then:- any challenge of the flow still pending → every pending
challenge of the flow is returned again,
reissued: true(not counted; SDK hooks do not run again; nothing is consumed) — whether its id was presented or withheld; - otherwise the flow's results are the flow state, and this verdict consumes every unconsumed challenge of the flow (a withheld one included): one set of passes buys one decision.
Without a usable presented id → step 5.
- unknown, other action or user → no credit (info reason
-
Open-challenge lookup (withholding the ids never helps): with no usable presented id, the latest unconsumed challenge of the same action and user (or, without a user, the same device), issued within the longest check timeout + 15 minutes and not passed, attaches its flow, with the same rules as step 4: pending challenges of the flow → re-returned (
reissued); otherwise the flow's results count and all its unconsumed challenges are consumed (info reasonchallenge_attached). The lookup never attaches a flow through a passed challenge. Token-less verdicts without a user have no binding key; the network counters (§5) cover them. -
App issue cap (SMS pumping): an app challenge is not issued when the check's
max_issued_per_hour(default 5) is reached for the user, else the device, else the IPv4 /24 · IPv6 /48 → the check'son_incomplete,decided_by.type = limit, codecheck_rate_limited,retry_after_seconds: 3600. The cap is per check: when an app check of a parallel group is capped, that check'son_incompletedecides and nothing of the group is issued.
Steps 4–6 run on every action (there is no action mode; only an enforcing
rule issues a challenge). The attached flow is locked as a whole — every
challenge of it, SELECT … FOR UPDATE in id order — so concurrent verdicts of
one flow serialize; consuming, issuing the next challenges and recording the
verdict happen in one transaction. A verdict that involves no challenge flow
(nothing presented, no open challenge, none to issue) is recorded behind its
answer instead (§8, invariant 10).
2.3 Rules (engine.Ruleset.Decide, a pure function)
matches = enabled rules in scope whose condition is true
enforcing = matches of rules in state Enforcing (rule.mode = enforce)
1. any enforcing deny → deny
2. a check failed in this flow → its on_fail (deny; review only for score-based
widgets); several failed → the strictest (deny over review)
3. walk enforcing rules top → bottom (allow lists first):
allow → allow
review → review
challenge(group) → evaluate the group (below); satisfied → continue
4. nothing decided → allow
A check of a group is satisfied when it passed in this flow or its pass is remembered (a remembered pass uses one skip). Unavailable: the check's circuit breaker is open or its secret is missing or unreadable. Issued counts every challenge the flow has issued so far.
any (one of these, random by weight):
an item satisfied → satisfied
an item incomplete / error in this flow → that check's on_incomplete / on_error
(never a switch to another check of the group)
candidates = available items with weight > 0,
else available weight-0 items (fallbacks, in group order)
no candidate → the first check's on_error (check_unavailable)
issued ≥ budget → review (challenge_limit)
otherwise → challenge: ask {any, candidates}
sequence (all, one after another), items in order:
satisfied items are skipped; for the first other item c:
c incomplete / error in this flow, or unavailable → c.on_incomplete / c.on_error
issued ≥ budget → review (challenge_limit)
otherwise → challenge: ask {sequence, [c]}
every item satisfied → satisfied
parallel (all at once):
missing = items not satisfied; none → satisfied
missing items incomplete / error / unavailable → the strictest of their on_incomplete /
on_error (deny over review); nothing issued
issued + len(missing) > budget → review (challenge_limit)
otherwise → challenge: ask {parallel, missing}
- The budget is per attempt, across rules, and planned up front. A flow may
issue at most budget verifications (
settings.max_challenges_per_flow, 1–5, default 3), shared by every rule that matches it; a check two rules ask for counts once (it passes once for both). Before a step asks anyone to verify,Decideadds up what the matching challenge rules from this one on still need (top to bottom, until an Allow / Your app decides rule ends the walk; a rule whose check already failed, was not completed, errored or can't run ends it too): sequence and all-at-once groups every unsatisfied check, "one of these" one — the fewest possible (a check an earlier rule plans counts as passed), so a flow that can fit is never ended early. If issued + need > budget, the flow ends now with review (challenge_limit) instead of after the person did the first checks; itsdecided_bynames the first rule that does not fit (id,name). The panel says which rule did not fit, and warns when rules together can need more than the budget. - Decide never picks. For a challenge it returns the ask: the
group's mode and the checks to issue now (any: the candidates with their
weights; sequence: the next check; parallel: every missing check). The
service then applies the app issue cap (§2.2 step 6) and issues: for
anyit picks one candidate at random by weight (crypto/rand); the other modes issue every asked check. The issued checks are stored on the challenges and in the verdict inputs (picked), soDecidestays pure and a replay reports what was issued. - Weights (
anyonly): integers 0–100, default 1, at least one above 0. Weight 0 = fallback only: when every weighted check is unavailable, the first available weight-0 check (in group order) is issued. The other modes carry no weights. - No re-rolling. A pending challenge is re-returned (
reissued), never drawn again, so retrying cannot pick an easier check; a failed, incomplete or erroredanychallenge applies that check's outcome. Each new flow (a new attempt) draws again. - A group's checks count toward the flow's 3 challenges and "each check once" (a check with a result in the flow is never issued again).
- Would-run: when a monitoring rule matched, the same walk over every
match (monitoring rules as if they enforced), without side effects →
would_decisionandwould_challenge({mode, checks}of a would-challenge — foranythe candidates;nullotherwise); the monitoring matches are listed as reasons withmonitor: trueand inmonitor_rules. - State of a rule: Off when disabled, else Enforcing or Monitoring by
its
mode— the same in every action of its scope. It is the only monitor/enforce switch (actions have none). - Rule health: the panel lists per rule its
unseen_attributes— theuser.*/context.*attributes the condition reads that were never received in traffic (the rule cannot match on them: a missing value compares false); an enabled one raises the to-dorule_cant_fire({rule, rule_id, attrs}). - To-dos for monitoring rules:
rule_ready({rule, rule_id, n, what}) for each monitoring rule with would-hits in the last 7 days (what: the would-decision, or the comma-separated check keys of a challenge rule);rules_monitoring({n, rules}) while enabled rules still monitor and the project has traffic;no_enforced_rule({actions}) for actions with traffic that no enforcing rule covers. - Remembered pass: widget checks — a pass on the same device id,
fingerprint hash and IPv4 /24 · IPv6 /48, with
device.users_24h ≤ 2; app checks — a pass of the same user (and device when both have one). Within the check'sremember_seconds, at most 20 skips per pass. Always the browser's own device id: a device recognition linked the browser to (§2.6) never lends its passes. Remembered passes are per check: in a group each satisfies its own item (any: one item's remembered pass satisfies the rule). decided_by:{type: builtin | token | rule | check | limit | default | test, id?, name?, code}— "Your app decides" is never ambiguous. Monitoring rules never decide.passed("sms")in a condition reads the same flow/remember state.
2.4 Conditions and fields
- expr-lang programs, at most 500 nodes and 2000 characters, 200 rules per project.
- Typed attributes are fields:
user.api_calls_total,user.registered_at,context.amount. They must be declared (observed in traffic, set on the Attributes page, or declared by a template on save); an undeclared one is a compile errorattribute_undeclared. Keys may not shadow built-in user fields (id,present,age_days,verdicts_1h,devices_30d,new_device,device_age_seconds,challenges_issued_1h). - Missing values: any comparison with a missing value is false
(
!=too);has(context.amount)tests presence;!(context.amount >= 100)is true when missing (the compiler warns). - A value of the wrong type is read as missing and counted on the attribute
(
wrong_type); retyping or deleting an attribute a rule uses → 409attribute_in_use. - Helpers:
days_since(ts),hours_since(ts),in_list(list, value),cidr(ip, prefix),header(name),now(),passed(check). - Events the server reports:
event_count(type, window),event_sum(type, window),event_last(type); the user's attempts at an action:attempts(action, window)(§5.1); totals across users:counter(key, window)(§5.2);email_domain(value). - Risk (a hint):
risk.scoreis the sum of the leading signals (within a group only the largest weight counts; clamped 0–100);risk.levelnames its band with the project's thresholdssettings.risk(default low < 25 ≤ medium < 50 ≤ high < 75 ≤ critical; 1 ≤ medium < high < critical ≤ 100),unknownwithout a valid token. Weights, on/off and thresholds are per project; they are compiled into the ruleset (a change bumps the rules version), so a change applies from the next verdict and never rewrites a stored verdict's level. Nothing decides on risk unless a rule reads it.
2.5 Baseline rules (seeded, editable, "recommended", monitoring)
Every new project gets these two rules in Monitoring: they report what
they would do (would_decision) and decide nothing until someone enforces
them. Once they have would-hits, the to-do rule_ready (one per rule, as
for any monitoring rule) suggests enforcing each once they look right.
| Rule | Outcome |
|---|---|
Repeated failed or unfinished verifications: device.present ? device.challenges_not_passed_1h >= 5 : network.challenges_not_passed_1h >= 50 |
Block |
High risk: risk.level in ["high", "critical"] |
Your app decides until a widget check exists; the to-do high_risk_without_check turns it into Verify with that check |
No default rule blocks on risk alone ("Critical risk → Block" is a strict template).
2.6 Recognised devices (privacy.device_id = recognition)
Opt-in per project. Recognition links, it never merges: a browser that
presents no valid device id (a private window, cleared storage) always gets
its own new device id (cookie and localStorage as usual). When its keyed
fingerprint print matches a device the project saw in the last 30 days, the
assessment only records a probabilistic link to that device
(assessments.recognized_device_id + a confidence; docs/DESIGN.md, "Device
recognition"). Nothing ever writes a known device id into a new browser.
It is deliberately narrow, because look-alike browsers are common:
- Desktop browsers only: Windows, macOS, Linux and ChromeOS, with at
least 3 of the canvas / WebGL / audio / fonts hashes. iPhone and iPad (any
browser: all are WebKit, one fingerprint per model and iOS version),
Android and other mobile or touch-first platforms are never recognised and
their prints are never stored (mobile User-Agent,
Sec-CH-UA-Mobile: ?1, a non-desktopSec-CH-UA-Platform, or touch points on a macOS / Linux User-Agent — iPadOS and Android's "desktop site"). - Same network only: the linked device was last seen on the request's IPv4 /24 · IPv6 /48 (both confidences), from a public address. No ASN matching (a carrier's ASN holds millions of users).
- Busy networks are skipped: a /24 · /48 on which more than 20 devices left a print in the last 7 days (CGNAT, offices, campuses, hotels) is never used.
- Shared prints are dropped: when two established device ids (each lived at least 24 hours) with the same print are in use at the same time (identical machines, two browser profiles on one computer), the print is marked shared for the project and never linked again — neither exactly nor as a fuzzy candidate. Identical machines on one network may be linked once, until the second one's id is established.
high— the same stable print as exactly one device (with the browsers already linked to it) on the network block;medium— most components equal, same network block, same browser family, and no equally good other device. Anything weaker, ambiguous or shared is not linked.- A presented id always wins (no recognition runs); a randomised probe
(
js.fingerprint_randomized, signalfingerprint_randomized: Safari private mode, Firefox with resistFingerprinting, Brave) is never recognised.
| field | type | meaning |
|---|---|---|
device.recognized |
bool | no device id was sent; recognition linked this browser (its own new id) to a known device |
device.recognition |
high · medium · "" |
confidence of the link |
device.recognized_blocked |
bool | the linked device (or a browser linked to it earlier) is on block_devices, or a user on block_users was seen with it |
device.recognized_users |
number | distinct users seen with the linked device (and the browsers linked to it) in the last 30 days |
js.fingerprint_randomized |
bool | canvas / audio fingerprints changed between identical renders (privacy browsers, private modes, anti-detect tools) |
A link is a hint for your rules, never an automatic decision and never
carried trust: every other device.* field (age, users, denies,
blocked-user flag, allow and block lists, challenge_passed_recently,
remembered passes) and user.new_device are the browser's own new id's. No
built-in deny comes from a link — the linked device's block list entry only
shows as device.recognized_blocked, which your rules can act on (e.g.
device.recognized && device.recognized_blocked → Verify). Recognition can
therefore never make an outcome worse than no recognition, except through
your own rules on these fields; a private window of a blocked device is a
new device, as without recognition. The fields are stored with the verdict
(device.recognized_blocked, device.recognized_users as counters) for
replay.
3. Checks
| Field | Meaning |
|---|---|
key |
slug rules and SDK handlers use (captcha, sms) |
kind, provider |
widget: turnstile · recaptcha · recaptcha_enterprise · hcaptcha · friendly_captcha · geetest · abusend_pow · abusend_hold · abusend_trace; app: sms · email · kyc · custom |
config |
widget only: site_key (GeeTest: the 32-hex captcha_id), recaptcha_version (v3 · v2_invisible · v2_checkbox), gcp_project_id, threshold, appearance, region (Friendly Captcha: global · eu); abusend's own checks take only pow, strictness and appearance (no site_key; pow / strictness on another provider → 400 invalid_check) |
config.appearance |
invisible (default, stored omitted: the widget runs in the background and only shows up when the provider wants an interaction) · visible (the browser SDK shows the widget — checkbox, puzzle — in its verification dialog). Turnstile, hCaptcha, Friendly Captcha, GeeTest: both; reCAPTCHA v2_checkbox, abusend_hold, abusend_trace: visible only (stored visible); reCAPTCHA v3 / v2_invisible, Enterprise and abusend_pow: invisible only. Anything else → 400 invalid_check |
config.pow |
own checks: the proof-of-work puzzle's size as work (2^work hashes expected; each step doubles it), {mode: risk | fixed | off, fixed, low, medium, high, critical}, every size 10–26; default risk with 15 / 18 / 20 / 22 (low / medium / high / critical), fixed 18. risk takes the size of the verdict's risk.level when the challenge is issued (unknown, e.g. no browser token → low); off (no puzzle) only for abusend_hold / abusend_trace |
config.strictness |
abusend_hold, abusend_trace: lenient · normal (default) · strict — the movement score an attempt needs: ≥ 0.3 · 0.5 · 0.7 (§3.1) |
| secret | widget only, AES-GCM, AAD bound to (project, check), never returned; abusend's own checks have none (a secret → 400 invalid_check) and are always configured |
remember_seconds |
0–604800, i.e. up to 7 days (default 24 h widget / 0 app); remember_hours = whole hours, rounded down (older clients; accepted when remember_seconds is absent) |
on_fail |
deny (review only for reCAPTCHA v3 / Enterprise; not for v2 invisible / checkbox) |
on_incomplete |
review (widget default: ad blockers hit real users) · deny (app default: a real user retries) |
on_error |
review · deny (never allow) |
timeout_seconds |
widget 30–300 (120), app 60–1800 (600) |
max_issued_per_hour |
app only, 1–100 (5) |
- A check a rule uses cannot be deleted (409
check_in_use). - hCaptcha and GeeTest have no action/nonce binding (GeeTest's answer has
no hostname either):
/verifyrequires the challenge's device cookie and an allowed origin (true of every challenge with a device). Friendly Captcha has no action/nonce either; its answer's origin is checked against the allowed origins. - GeeTest v4: the browser sends
provider_token= the compact JSON of the widget'sgetValidate()object (lot_number,captcha_output,pass_token,gen_time, strings only, nothing else); a malformed one fails the challenge without calling GeeTest. - App results: the customer reports
failedonly after its own retry limit (a mistyped code is not a failed check).
3.1 abusend's own checks
Three widget providers the edge verifies itself: no site key, no secret, no
provider script, nothing sent to a third party and no egress
(internal/owncheck). The browser SDK (3.3.0) runs them from the challenge's
params (§4) and answers /verify with a solution instead of a
provider_token.
| provider | appearance | what the browser does |
|---|---|---|
abusend_pow |
invisible | solves a proof-of-work puzzle in the background |
abusend_hold |
visible | in the verification dialog, presses and holds two dots in order until each ring fills (targets random per challenge, hold_ms 700–1200); the puzzle is solved meanwhile |
abusend_trace |
visible | moves the pointer or a finger around a 16:10 box until a bar fills (3–5 s of movement, random per challenge); the puzzle is solved meanwhile |
-
Puzzle (
params.pow:{alg: "sha256", seed, bits, puzzles, work}): for each i in [0,puzzles) (8) find a number n (decimal) so that SHA-256("<seed>.<i>.<n>") starts with at leastbitszero bits;bits=work− 3, so the 8 puzzles together cost 2^workhashes on average (several small puzzles keep the time near its average). The seed is the challenge's randomnonce. The size is fixed when the challenge is issued (config.pow, §3: by the verdict's risk level, fixed, or off for the visible checks) and stored with it; the answer is checked against it. -
Tasks (
params.task):{type: "hold", targets: [{x, y}, {x, y}], hold_ms}(centres as fractions of the box) or{type: "trace", duration_ms}. The SDK records the pointer events on the task's box only (solution.telemetry, §4) and the edge scores them. -
Score (0–1, stored as the challenge's
score): heuristics, not a model — what a person's movement usually looks like against what scripts usually produce. Each finding lowers the score by its weight (owncheck.Hints:untrusted_events(events the page dispatched itself),teleport,jittery_noise,linear_path,exact_center,constant_turning,constant_speed,jump,periodic_motion,too_fast,identical_holds,instant_reaction,instant_release,regular_timing,no_submovements,straight_path). An attempt passes when the task was done and the score reaches the check'sstrictness(0.3 / 0.5 / 0.7). -
Attempts (3 per challenge, like every widget): a visible check whose attempt did not pass while attempts remain answers
{status: "pending", retry: true, missing?}and stays pending; the SDK shows a short hint and starts the task over.missingsays what was not done:presses,target_missed,hold_short,trace_short,trace_small; it is omitted when the task was done but did not convince. -
Outcomes (
detail):outcome status · detail applies puzzle not solved failed · pow_invalid(no retry)on_failmalformed telemetry failed · rejectedon_failthe movement did not convince, after the last attempt failed · not_human_likeon_failthe task was not done, after the last attempt incomplete · task_incompleteon_incompletethe same recording again failed · token_reusedon_failno usable params(e.g. issued by an older pod)error · provider_erroron_errornot run (script blocked, timeout, closed) the SDK's errorreport, as for any widgeton_incomplete -
Replay: a recording's signature is a SHA-256 of its relative movements (1 px, 1 ms steps), so the same recording sent again — even moved elsewhere in the box — is caught like a reused provider token (the hash is kept in
challenge_tokensfor 1 day). Recordings under 10 events have no signature. -
Stored: the status, the detail, the score and the replay hash. The raw recording is never stored or logged; the panel's preview (below) stores nothing.
-
Limits: like every client-side check these can be beaten — a solved puzzle proves spent CPU, not a person, and a recorded or synthesised movement can score well. They add cost and friction for automation; the score is a hint the strictness turns into pass / fail. The visible tasks need a pointer or touch: keyboard-only visitors cannot do them, so pair them with another check (an
anygroup, e.g. with SMS) where those visitors must get through. -
The panel's test button answers ok for own checks (there is no provider to reach); its preview scores a recording made in the panel (
POST /projects/{p}/checks/preview,docs/PANEL.md§5.2).
4. Edge API
| Endpoint | Auth | Body → answer |
|---|---|---|
POST /v1/assess |
site key | signals (sealed probe) → {token, expires_in, device_id} |
POST /v1/verdict |
secret key | {token?, action (required), user: {id, attributes}, context, client_ip, challenge_ids?, review_handling?, simulate?} → verdict (challenge_ids: 0–3 ch_… ids) |
POST /v1/challenges/{id}/verify |
site key | widget: exactly one of {provider_token}, {solution} (abusend's own checks only) or {error: script_blocked | timeout | widget_error | abandoned} (pending only; never passes an app check; body ≤ 192 KB) → {status}, for an own visible check also retry, missing |
POST /v1/challenges/{id}/result |
secret key | {status: passed | failed | abandoned, user_id} → {status}; 409 challenge_resolved, challenge_user_mismatch, challenge_kind_mismatch |
GET/PATCH/DELETE /v1/users/{id} |
secret key | user attributes, erasure |
POST /v1/feedback |
secret key | {verdict_id, label: fraud | legit, note} (labels the whole flow) |
POST /v1/events |
secret key | {events: [{user_id, type, value?, properties?, occurred_at?, id?}]} (1–100, all or nothing) → {accepted, duplicates}; 400 invalid_events, invalid_event, invalid_occurred_at, event_value_out_of_range (§5.1) |
Verdict response:
{ "id": "…", "flow_id": "fl_…", "decision": "challenge", "would_decision": null, "would_challenge": null,
"decided_by": {"type": "rule", "id": "…", "name": "Dormant account, big first top-up", "code": "rule_…"},
"challenges": [{"id": "ch_…", "check": "sms", "name": "SMS", "kind": "app", "provider": "sms", "expires_in": 600, "reissued": false}],
"risk": {"score": 12, "level": "low"},
"reasons": [{"code": "rule_…", "label": "…", "kind": "rule", "effect": "challenge", "rule_id": "…"}, {"code": "ip_datacenter", "kind": "signal", "weight": 25, "label": "…"}],
"signals": {"ip": {"address": "…", "country": "SG", "asn": 7473, "datacenter": false, "vpn": false, "tor": false, "relay": false},
"client": {"kind": "browser", "label": "chrome", "ja4": "…", "ua_family": "chrome"}, "probe": "ok",
"device": {"present": true, "id": "d1.…", "new": false, "users": 1}},
"token_status": "valid", "user": {"id": "u_…", "attributes": {}}, "action": "topup", "test": false,
"created_at": "…", "retry_after_seconds": null }
challenges holds 1–3 challenge objects when the decision is challenge
(all to run now; several only for a parallel group) and is [] otherwise.
would_challenge is {mode, checks} for a would-decision challenge, else
null. Widget challenges add site_key, mode (turnstile, recaptcha_v3,
recaptcha_v2_invisible, recaptcha_v2_checkbox, recaptcha_enterprise,
hcaptcha, friendly_captcha, geetest_v4, abusend_pow, abusend_hold,
abusend_trace), action, nonce, script_url, timeout_ms, appearance,
ui and, for abusend's own checks, params:
appearanceis"visible"when the SDK shows the widget in its verification dialog, omitted when it is invisible. The mode stays the same for both appearances of Turnstile, hCaptcha, Friendly Captcha and GeeTest (SDKs that predateappearancekeep the invisible path);recaptcha_v2_checkboxis always visible.timeout_ms: invisible widgets at most 30 000; visible widgets wait for the person until the challenge expires (the check'stimeout_secondswhen issued).region(Friendly Captcha only) is"eu"when the check verifies with the EU API, so the widget uses it too (itsapiEndpoint: "eu"); omitted: global.params(own checks only, §3.1):{pow?: {alg: "sha256", seed, bits, puzzles, work}, task?: {type: "hold", targets: [{x, y}, {x, y}], hold_ms} | {type: "trace", duration_ms}}—powis absent when the check runs without a puzzle,taskforabusend_pow. Themodeis the provider name; there is nosite_keyorscript_url, andactionis not used.uiis the project'ssettings.challenge_ui—{"theme": "auto" | "light" | "dark", "accent": "#rrggbb" (omitted when unset), "branding": true | false}(defaultsauto, no accent,true: a small "Protected by abusend" line).
428 contract (what protect() / guard() answer and Abusend.fetch
understands, also for native/mobile clients): status 428, header
Abusend-Challenge: 1, body {"error": "challenge_required", "challenges": [{…}]}. The client runs every challenge and retries with a fresh token and
X-Abusend-Challenge: <id>,<id> (comma-separated, every id of the 428).
/v1/challenges/{id}/verify and /v1/challenges/{id}/result stay per
challenge.
/verify for abusend's own checks: the body carries solution =
{pow: [n, …] (one nonce per puzzle, [] without a puzzle), telemetry?};
telemetry (the visible checks) is {v: 1, w, h (the box in CSS px), dpr, pt: mouse | touch | pen, ev: [[t, x, y, type], …] (≤ 3000; t in ms since the box was shown, x / y in px inside it, type 0 move · 1 down · 2 up), ut (events the page dispatched itself)}. A solution for another provider, or a
provider_token for an own check → 400 invalid_request; more than one of
provider_token, solution, error → 400 invalid_request. The answer is
{status}; while a visible check has attempts left and the attempt did not
pass: {"status": "pending", "retry": true, "missing": "hold_short"}
(missing omitted when the task was done but did not convince).
5. Counters (fields)
From the challenges table, computed only when a compiled rule uses them:
device.challenges_not_passed_1h, network.challenges_not_passed_1h
(failed + incomplete), device.challenges_issued_1h,
network.challenges_issued_1h, user.challenges_issued_1h. The network is
the IPv4 /24 · IPv6 /48 (careful with shared networks: CGNAT, offices).
Velocity hints, not exact totals:
- Capped. A challenge counter counts no further than its cap: the
largest number any enabled rule compares it with, plus one, at least 100 and
at most 1000 (a counter a rule reads in any other way — arithmetic, compared
with another field, or through the bare
device/network/userobject — is capped at 1000). Every comparison with a number gives the same result as without the cap; only the value itself saturates (a>= 50rule stores at most 100 however many rows exist). The storedinputs.countershold the capped value, so replaying a stored verdict with the ruleset that was live reproduces its decision (§6); a simulation of a new rule whose threshold is above the stored value sees the saturated value and can under-report. - Cached per pod. A pod reuses a counter for
challenge.counter_cache_ttl(default 2 s, 0 = off; the key is project, dimension, key, kind and cap; concurrent lookups of one key share one query). A counter is therefore up to that long behind the challenges table, on top of the per-pod view described in §5.2. What decides about a challenge is never cached: the issue cap (invariant 4), the flow lock, consumption, remembered passes. - Device and IP statistics are capped and kept in Valkey windows.
device.assessments_1h,device.users_24h,device.users_total(also shown assignals.device.usersin the verdict answer),user.devices_30d,ip.assessments_1h(stored on the assessment asip_count_1h) andfingerprint.devices_1hsaturate at a cap computed like the challenge counters': the largest number an enabled rule compares the field with plus one, at least 100 (max(100, threshold + 1)) and at most 1000 (fingerprint.devices_1h: 5000, its row window); a field a rule reads in any other way keeps the ceiling, a field no rule reads is capped at 100 (it was 1000). Sosignals.device.usersandip_count_1hshow at most 100 unless a rule needs more. Every comparison with a number gives the same result as without the cap, the storedinputs.countershold the capped value (replay is exact; a simulation of a new rule with a higher threshold can under-report, and a rule added between an assessment and its verdict sees the IP count saturated at the cap of the ruleset live at assess time), and a shared counter (§5.2) whose expression readsdevice,fingerprint,ipordevices_30drestores every ceiling.fingerprint.devices_1hcounts the distinct devices that used the fingerprint in the last hour, by last seen, including the asking device (a window of at most 5000 devices; at least 1),ip.assessments_1hthe assessments of the public address in the last hour before this one (private, loopback and carrier-grade addresses are not counted),device.assessments_1hthe device's assessments in the last hour. These three anduser.verdicts_1handdevice.denies_24hbelow are sorted sets in Valkey written when the assessment / the verdict is recorded (DESIGN "Device statistics caps and the Valkey windows"); the rest of the device statistics (device.users_24h,users_total,user.devices_30d, …) are one PostgreSQL statement. Nothing is cached per pod. Replay stays exact through the stored inputs. device.denies_24his the number of the device's deny verdicts in the last 24 hours, at most 1000 (it used to count the denies among the device's latest 1000 verdicts, so a very busy device could hide an old deny).- Stored counters omit zeros.
verdicts.inputs.countersholds only the non-zero counters; replay reads a counter that is not stored as 0 (older rows that store zeros replay the same). user.verdicts_1his the size of the user's window of verdicts in the last hour (Valkey), at most 1000 (it used to be an uncapped count over the verdicts table: a threshold above 1000 can no longer match), read only when an enabled rule reads it; otherwise the verdict stores 0 for it, like the other counters here. Testing or simulating a new rule that reads it against older stored verdicts sees 0 unless a live rule read it then.- Visibility. Verdicts are logged behind the answer (invariant 10):
attempts()(§5.1), written to Valkey behind the answer too, sees a verdict a few milliseconds after its answer; the panel's lists (ClickHouse) about a second later.user.verdicts_1handdevice.denies_24hsee it as soon as its window write (behind the response too) lands, typically a millisecond. The device counters (device.assessments_1h,device.users_24h,fingerprint.devices_1h, …) are read at the verdict, not at/v1/assess; the IP count is taken at/v1/assess.
5.1 Events
What abusend cannot see — a completed top-up, payment, withdrawal — the
customer's server reports after it happened with POST /v1/events (secret
key; Node SDK abusend.events.track):
{events: [{user_id, type, value?, properties?, occurred_at?, id?}]}, 1–100 per call, all or nothing (one invalid event refuses the batch; the error names it:details.index,details.field).user_id: the opaque id of the verdicts (1–256 characters; an unknown id creates the user like a verdict does).type:^[a-z0-9_]{1,32}$.value: a finite number within ±1e15 (summed byevent_sum; absent adds 0; a larger one is refused withevent_value_out_of_range, so no total can overflow).properties: up to 32 scalar values, 4 KB (stored, not read by rules).occurred_at: RFC 3339, default the time of receipt; more than 5 minutes in the future or older than 30 days →invalid_occurred_at; a few minutes ahead (clock skew) is stored as received.id(1–128 printable ASCII, no spaces): an event whose id was already reported in the project is a duplicate (not stored again, counted induplicates), whatever its content. An id is remembered for 31 days (a day more than an event may be old), by the whole fleet at once; erasing the user forgets the ids of their events (the newest 50 000 ids per user are listed for that; an older one of a user who reported more still dedups until it expires, erasure just cannot find it: it is a hash, no personal data).- Limits: per secret key a call budget of
verdicts_per_secondcalls/s (its own bucket: reporting never uses the verdicts' budget) and an event budget ofverdicts_per_secondevents/s (burst 2×, at least one full batch); beyond → 429rate_limited, nothing stored. - Rule functions, per verdict user, over a window of the rule's choice:
minutes, hours or days,
"1m"to"30d"("5m","90m","2h","3d"; events withoccurred_atin (now − window, now]):event_count(type, window)(a number),event_sum(type, window)(sum ofvalue),event_last(type)(the newestoccurred_atwithin 30 days, or missing: comparisons with it are false; usedays_since/hours_since,has(event_last("topup"))). Arguments are quoted literals; an invalid type or window or a computed argument is a compile error (invalid_condition, with a readable message). Without a user in the verdict, counts and sums are 0 andevent_lastis missing. A window has one canonical name ("60m"is"1h","1d"is"24h","48h"is"2d"), the one values are stored under. Caps: rules see at most the newest 5 000 events per user and type within the 30 days (a count saturates there; the oldest fall out first) and at most the newest 1 000 attempts per user and action. Precision:event_count,event_lastandattemptsare exact (an event counts whennow − window < occurred_at ≤ now); anevent_sumhas the precision of a minute at the start of its window — it adds the events of the minutes fromnow − window(rounded down to the minute) up to the current one, so it can hold up to a minute of events more than the count of the same window (it is read from per-minute, hour and day sums, so its cost does not depend on how many events a user has). - Attempts —
attempts(action, window): the user's attempts at the action in the window, this one included when the verdict is of that action (attempts("topup", "5m") > 3blocks the fourth top-up in five minutes). An attempt is a flow, not a verdict: the retry after a verification belongs to the flow that asked for it and counts the same as the first request, so passing a check never tips a user over a limit. A flow counts from the moment its first verdict is decided (actionlike the verdict's,^[a-z0-9_.-]{1,32}$), whatever the decision (a verdict decided before the rules, such as a token deny, is a flow of its own); 0 without a user. It is written to the fast store as it is decided (batched, off the answer's path), so the next request of the user counts it on the same pod at once and on any other pod a Valkey call later; two requests decided at the same moment may not see each other. - Computed lazily: only what the enabled rules read — events in one
call per verdict (the count of every standard window
1h,24h,7d,30dof those types and of the windows the rules name, and the last time; sums only for the windows anevent_sumof the type names, so a verdict storessum:<type>:<window>for those and nothing else), attempts in one call once the verdict's flow is known. Both are read from the fast store (Valkey), exact across pods; if it cannot be read the verdict fails (5xx, which the SDKs turn into review): a rule never reads a silent 0. The values are stored in the verdict'sinputs.events(count:<type>:<window>,sum:<type>:<window>,last:<type>in unix seconds,attempts:<action>:<window>), so replays reproduce them (§6); a value a stored verdict lacks (a new rule or window) is computed as of the verdict's time for replay and simulation. - Event types are discovered like attributes (sampled; at most 200 listed
per project, with first / last seen and an optional label). A rule reading
a type no event was ever received of lists
event.<type>inunseen_attributesand raises the to-dorule_cant_fire. - Retention: events are kept per
privacy.event_retention_days(default 90, 1–1095); rules read at most the last 30 days (the longest window). What rules read lives in Valkey for the project's retention, never longer than rules can use: events (and their sums) formin(event_retention_days, 31)days, attempts formin(retention_days, 30)days. A shorter setting therefore also shortens the windows: withevent_retention_days = 1,event_count("topup", "7d")sees one day; lowering it takes effect on the next verdict (reads are bounded by it) and the next write (trimmed), so a rule never reads an event the analytics store has deleted, and a live verdict agrees with its replay. Erasure (DELETE /v1/users/{id}, panel erase) deletes the user's events, and what rules read of them and of the user's attempts (in Valkey: with the erasure and again with its repeat, §8 invariant 10).
5.2 Shared counters
Everything above is one user's (or one device's, one network's). A counter adds up across users: every attempt at an action, or every reported event of a type, over a recent window, optionally per group.
- Definition (panel → Counters, MCP
save_counter, config as code):key(^[a-z][a-z0-9_]{0,40}$, what rules name it by),name,sourceverdicts(the attempts at actionof: the first verdict of each flow, counted after its decision whatever it is; the retry after a verification is not counted again; verdicts decided by the token built-ins are not counted) orevents(the reported events of typeof, each in itsoccurred_atminute; events older than 24 h are left out),filter(a condition in rule syntax choosing what is counted; empty: everything),measurecount(1 each) orsum(ofvalue, an expression such ascontext.amount, for verdicts; of the event'svaluefor events),group_by(an expression; each value has its own total:ip.country,email_domain(user.email); empty: one total),enabled. Event counters knowuser.*only (an event carries no request:ip.*,context.*… are refused). Filter, value and group_by are at most 1000 characters and cannot read event, attempt or counter functions; at most 50 counters per project. A group is text (numbers and true / false printed), at most 64 bytes. An amount (sum) must be a finite number within ±1e15: a larger one (or NaN, infinity) adds nothing and is counted inabusend_counter_dropped_total, so a total is always a finite number (the reportedvalueof an event is refused at the door withevent_value_out_of_range). - At most 10 000 live groups per counter (fleet-wide; a group is live 25
hours after its last write). A value of a new group beyond that is added to
one shared group named
__other__instead (counted inabusend_counter_groups_overflow_total; a real group of that name merges with it): an end user who controls agroup_by(an IP, an id, free text) cannot create unbounded state. A group already counted keeps counting for itself; group by something coarse (a country, a domain) and watch the metric. - Rules read
counter(key, window)(window"1m"to"24h"): the total of the request's group (group_by evaluated on the verdict), or 0 for an unknown group, a disabled counter or one that no longer compiles. A rule naming a counter that does not exist is refused (invalid_condition); deleting a counter a rule reads is refused (counter_in_use). The filter only chooses what is counted: repeat it in the rule when only those requests should be affected —ip.country == "SG" && counter("payments_by_country", "1h") > 100→ Verify with SMS;user.email contains "hotmail" && counter("hotmail_purchases", "15m") > 50000→ Verify with a captcha. - Exact across pods (no flush lag): every counted value is added, as it
is counted, to minute buckets in Valkey (kept 25 hours after the last
write) with hour sub-totals, so a 24 h window is read as at most 24 hour
sub-totals plus the minutes at its two ends (a cost that does not depend on
the traffic); a rule reads the minutes of its window (the window's first minute
whole, so up to a minute more), so what any pod counted a moment ago is
read by the next verdict of any pod. What can still be missing: values a
pod could not write (Valkey error, a full queue: counted in
abusend_counter_dropped_total). A verdict never counts itself (it is counted after its decision). If the totals cannot be read the verdict fails (5xx, review at the SDK), never a silent 0. The panel's totals for a counter (and Watch's) come from ClickHouse, a few seconds behind. - A counter counts from when it is created; changing what it counts
(source, of, filter, measure, value or group_by) starts it from zero.
Buckets are kept 25 hours. The value a verdict read is stored in
inputs.events(counter:<key>:<window>), so replays reproduce it; a counter a stored verdict did not read counts as 0 in replays, rule tests and simulations (totals are not rebuilt for the past). - Helper for groups and conditions:
email_domain(value)— the lower-cased part after the last@(""when the value is not an address).
6. Replay and simulation
- Verdicts store their inputs (the user's attributes as used, the
context, verdict-time counters, the flow state, and
picked: the checks the verdict issued). Re-evaluating a stored verdict with the rules that were live then reproduces its decision (now= verdict time). Verdicts recorded before the action mode was removed on a monitoring action (decided_by.type = monitor) replay as their rules' states say, not asallow. A challenge reports the stored pick when it is among what the rules ask for now; a new or changedanyrule is reported as "one of" its candidates. - Replays, rule tests, simulations and the risk preview read the ClickHouse
log: the verdicts (with their inputs, decoded from the stored JSON with the
decoders every reader uses, so a replay sees byte for byte what the verdict
stored), the assessments they name and the events and attempts of the next
bullet. The log is about a second behind the request path (the writer's
batches,
clickhouse.flush_interval): a verdict made a moment ago may not be in the sample yet. The sample is the verdicts stored withfirst_of_flow(the verdict that opened its flow), newest first, counted per flow. - Event values (§5.1) come from the verdict's inputs. A rule reading an
event type the verdict did not store (the rule did not exist then) gets
it computed from the events stored now, as of the verdict's time (events
reported later with an earlier
occurred_atcount too: an estimate) — in replays, rule tests and simulations only; stored verdicts never change. Attempts are computed the same way; a shared counter (§5.2) a verdict did not read counts as 0. So doesuser.verdicts_1hwhen no live rule read it (§5), and a challenge counter or device statistic is the capped value the verdict stored (§5): a rule tested against it with a threshold above that cap sees the saturated value. - Risk is recomputed in a replay from the verdict's assessment with the
replaying ruleset's weights and thresholds; stored verdicts keep the
risk_score/risk_levelthey were given.POST /projects/{p}/risk/previewrescores the most recent ≤ 5000 verdicts with a known level from their stored signal reasons with the current and the proposed weights and thresholds, and reports the level changes and, for every enabled rule readingrisk.level/risk.score, how many of them its condition matches now and with the proposal (its other fields from the replayed environment). Signals disabled when a verdict was made are not in its reasons and cannot be counted. POST /projects/{p}/rules/simulate {rule | delete, days}replays the first verdicts of flows of the lastdays(default 7, newest 20 000) as if the candidate enforced, with the current rules (before) and with the change (after), against current lists and config. It reports matched, before / after decisions, changed, per check the verifications and users per day who would see it (a check that is not connected is listed withconnected: false), and per attribute the share of verdicts that did not carry it. Ananygroup whose pick is not known shares its verifications by weight (expected counts). Items carrychallenge({mode, checks}). There is no minimum sample;tiersays how to read the report:none(no verdicts in the window;last_inputscarries the latest older verdict's action, user id, attributes and context, if any),small(belowsmall_below= 200: counts, not shares; every matching or changed flow is initems, each with its stored decision and risk level) orfull. Test-project traffic (forced decisions included) is replayed like any other. Challenge rules are replayed up to their first step (first_step_only).POST /projects/{p}/rules/testevaluates the rules against a stored verdict (verdict_id, as it was) or a made-up request (input: {action, user_id, user, context, token?}— no browser token, so no signals, and no counters; nothing is stored), optionally with a candidateruleas if it enforced (the answer then carriesbefore, the decision without it, andmatched). A challenge is reported aschallenge: {mode, checks}(also inbefore). Values that do not fit their attribute's type are listed inwrong_typesand read as missing, as in a verdict.- Rule hits are an hourly rollup (per-pod counters flushed every minute;
approximate); exact counts come from
verdicts.matched_rules/monitor_rules.
7. SDK behaviour
- Browser (
Abusend.fetch, 3.x): sends the token; on a 428 runs every challenge of it concurrently (widgets from the pinned provider hosts; app checks through the handlers registered withAbusend.onChallenge(check, handler)); when every one completed, retries with a fresh token andX-Abusend-Challenge: <id>,<id>; up to 3 rounds. Give-up (a handler rejects, no handler, a widget cannot load) aborts the other app handlers of the round, reports each as not completed and returns the 428 response. A 428 with more than 3 challenges or duplicate ids is returned untouched. Never rejects by default. - Forms (
data-abusend-action) are submitted throughAbusend.fetch; the final response is followed (redirected) or replaces the document;data-abusend-nativeopts out. - Node (
guard(Request)/ Expressprotect(), 3.x): forwards the header's ids aschallenge_ids(well-formed, de-duplicated, at most 3; optionchallengeIds), sendsreview_handling, runschecks.<key>(challenge, req)for each challenge of the verdict whosereissuedis false (a failing hook reports the verdict's app challengesabandonedand answers 503verification_unavailable), answers 428 (challenges) / 403 / the review handler.onReviewis required:"continue"continues only reviews decided by a rule or the token requirement; reviews decided by a check or a limit are answered 403verification_incompleteunlessonReviewis a function. Outages and 429 → review, never allow.events.track/trackManynever throw by default ({ok: false, error}andonError); events without anidget a random one so the SDK's retries are stored once.
8. Invariants
- Nothing fails open to
allow: outages, 5xx (503 busyand503 state_unavailableincluded: a token that cannot be redeemed is never treated as valid), 429 → review (or deny);on_fail/on_incomplete/on_error/token_missingare review or deny; the SDKs'onReview: "continue"never continues a review decided by a verification or a limit. An own check whoseparamsare missing or unreadable is an error (on_error), never a pass. - Denies always win: built-in token and list denies (whatever the rules' states) and enforcing deny rules beat any passed check and any allow rule.
- Withholding never helps: a caller that doesn't run a check, doesn't send
a token, omits or forges a challenge id, or retries after a failure
never does better than one that ran the check and didn't complete it.
With groups: completing one of two parallel checks never satisfies the
rule; withholding the other's id re-returns it while pending (and its
outcome applies once it has one); retrying never re-draws
an
anypick and a failed pick never switches to another check; a flow's results are consumed once (every unconsumed challenge of it, withheld ones included). A browser that leaves its device id behind is a new device; recognition only links it (§2.6) and never lends it a known device's trust. - App checks cannot pump SMS/KYC cost: re-returned challenges don't re-run hooks; issuing is capped per user / device / network.
- Challenges are bound to (project, action, user, device when issued with a
token), consumed once (5 s replay for the same caller), at most the
project's budget per flow (
max_challenges_per_flow, 1–5, default 3) counting every issued challenge (a parallel group's included), each check once; the app issue cap applies per check; app results need the secret key and the matching user. Decideis pure; replay of a stored verdict with its rules (and the signal weights and risk thresholds compiled with them) reproduces it. The random pick of ananygroup is the service's, made when the challenge is issued and stored (picked) for replay.- A passed widget check may outweigh a failed probe (the probe is a hint). Device recognition is a hint too: it links, never merges; a link never denies by itself, never carries trust and never changes the browser's own device id (§2.6).
- Test-project features (
simulate, forced decisions, provider test keys, localhost origins) never work on live projects. - Monitor is safe: new rules — the seeded baseline rules included, and rules drafted by templates, the assistant or MCP — start in Monitoring; the panel's rule editor saves as Monitoring unless Enforcing is picked. Enforcing a rule is always an explicit, confirmed step; a replay is offered in that confirmation but never required. A rule's state is the only monitor/enforce switch: newly discovered actions are labels and never change what applies. Watch (§9) creates its rules the same way and can do nothing more: it never enforces one.
- The verdict log is written behind the answer (
log.async, default on). A verdict that involves no challenge flow is answered before its row is written; the row (ClickHouse'sverdictstable, the only store of the verdict log), the end user's row change and the device link (PostgreSQL) are written by the pod's log writer within milliseconds (log.flush_interval, 10 ms; ClickHouse's own batch adds up toclickhouse.flush_interval, 1 s). Nothing that decides, consumes, issues, redeems or remembers is written this way: challenge issue, consumption, replay, results, token redemption (an atomic script on the token's record in Valkey), remembered passes, issue caps, erasure and every panel write stay synchronous and exact, and a request that presents challenge ids, finds an open challenge or has one to issue always takes the synchronous flow transaction (invariant 3 is unchanged; its verdict row goes to ClickHouse after the transaction commits). A pod that crashes loses the rows still queued (normally the last 10–20 ms of its plain verdicts, and what ClickHouse's writer had not inserted yet, about a second); a log queue that stays full forlog.enqueue_waitanswers503 busy(an outage for the SDKs: the failure decision, never allow, invariant 1); ClickHouse's writer queue that stays full for the same time drops the log rows, counted (abusend_ch_dropped_total), without delaying the answer further. What reads the log (the panel's lists, replays) lags by the log's delay;attempts(),user.verdicts_1hand the counters are Valkey state written behind the answer as well. The answer never depends on the write, and the storedinputsare the ones the decision used (invariant 6). Erasure fences the log and ClickHouse's writer: rows of the erased user still queued on any pod are anonymised before they are written, every pod is told, and the erasure is repeated about 30 seconds later. What a queued verdict names (its assessment, the challenges of its flow, its device's recognition prints, the hot record in Valkey) is erased with the user although the masked verdict no longer names it. "Erased" means, precisely: the erasure is recorded (apending_erasuresrow naming the user only by derived ids, never the external id) and its PostgreSQL part is done when the call answers200; ClickHouse and Valkey are attempted in the same call, and whatever fails or is still running is retried by a leader job with backoff until every step succeeded (AbusendErasurePendingalerts after 15 minutes). The API answer does not wait for it; completion is guaranteed, normally within a minute. The write-behind queues of other pods and Valkey writes still in flight on them are fenced by the erase notification and swept by the repeat.log.async = falsewrites the users and device links of every verdict synchronously.
9. Watch ("Nöbet")
Watch is continuous monitoring that notices abuse within minutes, lets the
organization's own assistant (bring your own LLM key, the same one as the
panel's assistant) investigate headlessly, and reports with proposals. It is
off by default (settings.watch.enabled), per project. It never changes a
live decision on its own: the only things it may do alone are add a
Monitoring rule and create a counter that does not exist yet. The design
(buckets, scanner, runs) is in docs/DESIGN.md "Watch"; the panel page in
docs/PANEL.md §3.7.
9.1 What it does
-
Traffic buckets. Every project's traffic is counted in 5-minute buckets per action (
traffic_buckets): attempts (the first verdict of each flow, after its decision; verdicts decided by the token built-ins count too), how many were denied and challenged, and how many verifications failed or were left unfinished — in total and per country and per ASN of the client IP (a failed verification under the country and ASN the verdict that issued it saw, stored on the challenge). It is recorded whether or not watch is on, so switching it on starts with history. Buckets are kept 8 days. The scanner and the counter checks read them from ClickHouse rollups (summed per key), so a scan sees traffic a few seconds after the pods flushed it, like the shared counters. -
The scanner looks at each project with watch on about every five minutes (an atomic claim per project: two pods never scan it at once). No LLM is involved. The windows are the last 5–10 minutes (the previous bucket and the one in progress), 15 and 30 minutes; the bucket in progress counts partially, and every window is compared as a rate per minute over its real length. The baseline is the median of the same slot on the previous 7 days (needs at least 3 days with traffic recorded near that time); without it, the average of the two hours before the windows (needs at least 10 attempts, so a project that has just started has no "usual" yet and raises nothing). Detectors, all pure functions over the bucket series:
Detector Fires when (sensitivity normal) volume_spikeattempts in a window ≥ 50 and its rate ≥ 3× the baseline (10× for the last 5–10 minutes); a quiet action with a zero baseline fires at the minimum dominant_country,dominant_asna key has ≥ 30 attempts and ≥ 30% of the window, and had < 5% in the baseline (unknown country / ASN 0 never; at most 3 per dimension and action) verif_failure_rate≥ 20 challenges in the window, failure rate ≥ 30%, at least 25 points and 2× above the baseline (no baseline: ≥ 60%) deny_rate,challenge_rate≥ 50 attempts, the share rose by ≥ 20 points and 2×, or a share that was ≥ 15% fell to a quarter (a rule or an integration may have broken) counter_near_thresholda counter group is at ≥ 80% of the literal threshold an enforcing rule compares it with ( counter("k", "1h") > 100) and below it: a heads-up before the rule fires. Judged per group (the 50 largest groups of a grouped counter; at most three findings per comparison); the text and the brief name the counter, window, the dimension (an IP address, a user, …), value, the threshold and the rule — the threshold is not a "usual" value, and there are no attempts or baseline. While the same comparison already has a rule-applying episode open (below), none of its groups raises this finding and open ones close silently. Thresholds that are computed, or written on the left of the comparison, are not readA counter group at or over the threshold is not an anomaly: the rule is doing its job. It never wakes the assistant, is never escalated and is not in the daily digest. Watch tells you about it with short notes (a run record with trigger
rule_activefor the first note and the digests,resolvedfor the last, no investigation), and it does so per rule comparison, not per group:- One episode per comparison. The comparison is the counter, its window and the literal threshold of the enforcing rule (the first rule that reads the same comparison wins; a rule with two comparisons gets two episodes). All groups that are over the threshold at the same time (all the IP addresses of a burst, say) belong to one episode, so a burst of 200 IP addresses is one note, one run record and one e-mail, not 200.
- The operator counts.
>and>=are told apart: with> 20a value of 20 is not over the threshold, with>= 20it is. - What is read. The 50 largest groups of the counter over the window
(
MaxCounterGroups). When all 50 are over the threshold the list may be cut off, and one more query counts the groups at or over the threshold exactly, so the number in a note (group_count) is exact, while the examples (top_groups, highest first) are at most 50 in the run record and at most 5 in the note itself. A group that falls out of the 50 largest while it is still over the threshold is not "gone": the episode only ends when no group is over the threshold for two scans. - Kinds.
started(the first note: the rule now applies to N values, the highest, up to five examples),grew(a digest: the number went up or there are new values, with how many there were at the last note and how many are new) andended(two scans in a row with no group over the threshold; the number is the most it applied to at once). There is at most one note per episode per scan. - Cooldown. After a note, the next
grewdigest is due when the number of values grew or a listed value is new and the cooldown has passed: 30 minutes after the first note, then 1 hour, 2 hours, 4 hours, … (doubling with every note sent, at most once a day). Nothing is sent while nothing changed, however often the scanner looks.endedis not delayed by the cooldown. - At most 6 notes an hour per project (started, grew and ended together;
the count is derived from the run records, so it survives a restart). A
note over the limit is held: it is recorded as a run with status
skippedand the errornote_limit, nothing is mailed, and its cooldown starts again (a held start is sent later asstarted). The first note that is sent afterwards says how many notes were held. Notes count towards the five alerts per scan (§9.3). - An
endednote whose start was never sent (it was held by the hourly limit) is not mailed: it is recorded as a run with statusskippedand the errorstart_not_sent, and it uses none of the hourly limit. - Wording. Every note names what the counter is grouped by in words
(
ip.address→ "IP addresses",user.id→ "users",device.id→ "devices",ip.country,ip.asn,network.prefix,action,client.ja4; any other expression → "values"), says what the rule does and that it is Enforcing, says whether anyone is blocked (a challenge or "your app decides" rule blocks nobody by itself), describes the counter in words, and gives the highest value and up to five examples. It says that it is a note, not an alarm, and what to do if the traffic is expected (raise the threshold or narrow the rule). - Rolling update. During a rollout an old pod may still send the old per-group notes once (at most five e-mails and the "N more" one); the entries of the old form are merged into one episode afterwards and are not announced again.
The approaching finding of the same comparison is dropped silently when a group crosses (no "back to normal" note), and none starts again while the episode is open. Sensitivity scales the 80%.
Sensitivity low (5× / 100 attempts / 45% …) and high (2× / 25 attempts / 20% …) scale the thresholds.
-
Wake-up policy. Each anomaly has a fingerprint (project, detector, action, dimension, key). A new anomaly sends an alert and starts an investigation; one that has grown ≥ 2× since its last report does the same (escalation); one still ongoing gets a short status update at most every 30 minutes (a notification, no investigation); one gone for two scans gets a "back to normal" note (a notification). Counter groups that reach their rule's threshold are notes, one episode per rule comparison (started, digests on a growing cooldown, ended), not anomalies (§9.1 table). Several findings of one scan share one investigation. Optionally a daily digest runs at a chosen UTC hour, and Run now starts one by hand.
-
The investigation is a normal assistant conversation owned by the watch owner: the admin who switched watch on (
owner_user_id). The first message is a generated brief (the findings with their numbers, what the agent may do alone, the report format); the system prompt gets a watch addendum. The tools act through the same in-process invoker as the assistant, with the owner's permissions capped at admin, and every change is audited asvia: watch. Report and pending proposals appear in the owner's assistant drawer; proposals are approved with the assistant's normal confirmation.
9.2 What it may do alone, and what needs a person
| Tool | In a watch run |
|---|---|
read-only tools (get_stats, search_verdicts, list_networks, simulate_rule, …) |
always |
create_rule |
alone, always saved Monitoring (the panel API keeps a rule created through watch Monitoring whatever it was asked), marked origin: watch; at most 5 per project per day by watch; only in the run's own project |
save_counter |
alone, only when no counter with that key exists (changing one would reset its count) |
everything else that changes data: set_rule_mode, update_rule, set_rule_enabled, delete_rule, add_to_list, remove_from_list, set_action_token, set_attribute_type, set_risk_thresholds, apply_todo_fix, and any tool added later |
never alone: the call waits as a pending proposal in the conversation until the owner approves it |
The policy reads only the tool's name and its parsed arguments. Nothing the model writes — or anything that came from traffic, such as an attribute, a user agent or an event property — can approve a call; the model is told that everything from traffic is untrusted data. Rules and counters watch created are listed in the run's record.
9.3 Limits and failure
- Budget per project:
max_runs_per_hour(default 4, 1–20) andmax_runs_per_day(default 24, 1–200), the organization'smax_steps, and a wall-clock limit of 3 minutes per run. A run the budget refuses is recorded as skipped, its alert still goes out, and the findings get their investigation once the budget allows. - Without a usable assistant (not enabled, no key, key unreadable) watch
sends the deterministic alert and records the run as
no_llm. If the owner is no longer an admin or owner of the organization the run is skipped (alerts still go out) and the to-dowatch_owner_missingappears; saving the watch settings makes the saver the owner. - Switching watch off is the kill switch: scans stop and the open findings are forgotten; traffic keeps being recorded.
- Alerts are not limited by the run budget, but at most 5 are sent per scan
(the rest are summarised in one "N more findings" mail; notes about rules that
apply, §9.1, count towards the five), and rule-applying notes are limited to
6 an hour per project: the ones over the limit are recorded as skipped
runs with the error
note_limitand counted in the next note that is sent. - Notifications:
watch_alertandwatch_report(e-mail to the chosen org members, and the webhook eventswatch.alertandwatch.report; payloads in INTEGRATION.md). They carry numbers, the action, a country or ASN and the model's summary. A note about a rule that applies to the values of a counter also carries up to 5 of those values as examples (the run record keeps up to 50), and they can be IP addresses, user ids or device ids when the counter is grouped by them; nothing else about the end user. - The customer's LLM provider receives traffic excerpts during a run (PRIVACY.md).