api reference · v1

POST /v1/predict

Updated May 12, 2026

Classify a single piece of text. Returns a calibrated spam score between 0 and 1 — higher means more likely to be spam — up to a published ceiling, at which the score stops being a probability and becomes a lower bound. Read Confidence ceiling before you compare it to anything, and reading the score for the four comparisons that follow from it.

Data handling during the paid shadow trial

Normally, prediction text is processed in memory and discarded. When Siftfy enables the paid TypeSafe shadow trial, selected Pro requests with a local score from 0.15 to 0.85 and at most 20,000 UTF-8 bytes of text are retained for 30 days in restricted evaluation storage, with credentials removed, and sent to TypeSafe for background comparison. Free requests are excluded. The response below remains the local model response; TypeSafe does not change it.

Read the retention and provider-processing disclosure. Evaluation text is not used for training.

Request

Authenticate with the X-API-Key header. The body is a single JSON object:

fieldtypenotes
textstring1–20,000 characters; must contain at least one character that is neither whitespace nor a zero-width format character. Truncated to 512 tokens server-side.

Both checks run after Unicode normalization: a body that is empty, whitespace-only, or made only of zero-width format characters (Unicode category Cf) is rejected with 422, and so is a body whose normalized form is longer than 20,000 characters even when the raw string is shorter.

curl
curl -sX POST https://api.siftfy.io/v1/predict \
  -H "X-API-Key: $SIFTFY_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "Win a free iPhone! Click here NOW"}'
python
import os
import requests

resp = requests.post(
    "https://api.siftfy.io/v1/predict",
    headers={"X-API-Key": os.environ["SIFTFY_KEY"]},
    json={"text": "Win a free iPhone! Click here NOW"},
    timeout=5,
)
print(resp.json()["spam_probability"])  # 0.65
javascript
const resp = await fetch("https://api.siftfy.io/v1/predict", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "X-API-Key": process.env.SIFTFY_KEY,
  },
  body: JSON.stringify({ text: "Win a free iPhone!" }),
});
const { spam_probability } = await resp.json();

Response

json
{
  "spam_probability": 0.65,
  "likelihood": "medium",
  "detected_language": "en",
  "language_confidence": 1.0,
  "model_version": "bert-tiny-spam-signals+cal3",
  "max_confidence": 0.84,
  "score_source": "model"
}
fieldtypenotes
spam_probabilitynumberFloat in [0, 1]. Scores at or above 0.50 enter the review band. Equal to max_confidence it is censored — see below.
likelihood"low" | "medium" | "high"Coarse bucket for display and logging: low (<0.50, clean), medium (0.50–0.85, queue for review), high (≥0.85, reachable only with trusted feedback evidence). Not a routing field — high starts above the ceiling, and medium covers the graded band and the censored ceiling together. See reading the score.
detected_languagestringDetected language code used for calibration.
language_confidencenumberConfidence in the detected language, in [0, 1].
model_versionstringTraceable model and calibration revision.
max_confidencenumberHighest score the model and untrusted content signals may assert. If your block threshold is higher, model-only blocking is unreachable. It is also the value at which spam_probability stops carrying information — see below.
score_source"model" | "typesafe"Who produced spam_probability. model is our model, read as described on this page. typesafe means the key is paid, live escalation is on, our model was unsure (0.50 up to max_confidence), and TypeSafe answered with confidence, so spam_probability is TypeSafe's. Added 2026-09-23; every other field is unchanged. See TypeSafe scores.

Calibration

Scores combine the calibrated classifier output with bounded content evidence. Model and lexical evidence are capped below the 0.85 block band; only a confirmed feedback correction can exceed that ceiling. Pick a review threshold appropriate to your false-positive tolerance and validate it against your own traffic.

Calibration is fitted on short user-generated and form text — comments, contact and signup forms, account requests. That is the population the score is meaningful over. Pointing the endpoint at a different one, such as inbound email, is a reasonable thing to do and it is your decision to make, not an assumption you can inherit from this page.

Confidence ceiling

Every response carries max_confidence: the highest score the model is permitted to assert on its own word, so that it can never auto-block without a confirmed human report. The score is clamped to it. That clamp has a consequence for the number that is easy to miss and expensive to miss.

A score equal to max_confidence is censored. It means “at least this”, not “this”. Text the model reads as 0.86 and text it reads as 0.9999 both arrive as the same value. Above the clamp the score does not rank: a 419 letter and an invoice reach you identical, and the distance from that value to your threshold is not a measurement of anything.

So the score at the ceiling does not support arithmetic. Subtracting a discount from it, blending it with your own signals, or weighting it are all operations on a lower bound, and they produce a confident-looking number with nothing underneath it:

javascript
// WRONG. At the ceiling the score is censored, so this subtracts from a
// number that is only a lower bound. Whatever you subtract, every clamped
// message yields the SAME answer -- ceiling minus credit -- whether the text
// truly read 0.86 or 0.9999. So if that value lands just under your
// threshold, it lands there for all of them, every time. The arithmetic is
// perfectly deterministic. It is just meaningless.
const blended = result.spam_probability - myAuthCredit;
if (blended >= myThreshold) warn();

// WRONG for a second reason: 0.84 is not yours to hold. The ceiling is
// per-language and rises when a better checkpoint lands, so a copy in your
// code goes stale without failing.
if (result.spam_probability === 0.84) saturated();

// RIGHT. Three bands, read off the response. Compare; never subtract.
const { spam_probability: score, max_confidence: ceiling, score_source } = result;

if (score_source === "typesafe") {
  // Paid keys with live escalation: the model was unsure and TypeSafe
  // answered with confidence. This score is TypeSafe's spam probability, not
  // a clamped model score, so the bands below do not apply to it. A value
  // above the ceiling here is TypeSafe's call, not a human confirmation.
  score >= 0.5 ? block() : allow();
} else if (score > ceiling) {
  // ABOVE the ceiling. Not the model's opinion at all -- the only thing that
  // can exceed it is a human-confirmed spam report on this exact content.
  // Stronger evidence than anything below, and safe to act on.
  block();
} else if (score >= ceiling) {
  // AT the ceiling: censored. The model is as confident as it is permitted to
  // be on its own word, and the true value could be anywhere above this.
  // Unranked and suspicious -- queue it, or let your own signals decide. Do
  // not read the distance to your threshold as if it measured something.
  review();
} else if (score >= 0.50) {
  // BELOW the ceiling: a graded score. This is the band where the number
  // ranks, and the only band where arithmetic on it means anything.
  review();
}

Note that the comparison sorts responses into three bands, not two. Below the ceiling the score is graded and ranks. At the ceiling it is censored. Above the ceiling, when score_source is model, it is not the model's opinion at all — nothing but a human-confirmed spam report on that exact content can exceed the clamp, which makes it the strongest signal we return and the one case where a high score means what it appears to mean. Collapsing the last two into one branch throws that away. When score_source is typesafe, a score above the ceiling is TypeSafe's probability instead, and the three bands do not apply.

TypeSafe scores on paid keys. Since 2026-09-23, when live escalation is on for a paid key, a local score from 0.50 up to and including max_confidence is sent to TypeSafe while the request waits, for up to one second. A confident answer replaces spam_probability with TypeSafe's spam probability and sets score_source to typesafe; a threshold you already apply to spam_probability catches it with no code change. With no confident answer, or once the day's TypeSafe budget is spent, the local score comes back unchanged with score_source: "model" and HTTP 200. Free keys are never escalated. See the live escalation disclosure.

The ceiling is not a rare state. How often it fires depends entirely on your population, and on some populations it is the common case. The ceiling exists because of our own measurement: with the clamp off for English, two production sites enforcing at the block band refused legitimate submissions — 16 of 38 ordinary transactional requests in one four-hour window, and 3 of 3 in another measured set. Re-scored, those false positives came back at 0.9947–1.0000, where genuine spam sits, and near-identical short sentences score 0.0010 and 0.9990. On that kind of text the model has no separating signal, so almost everything confident clamps.

What the model can still do is rank within a corpus, which is what the review band is for. Build on the ranking below the ceiling and on your own evidence at it — not on the clamped value, and not on an assumption that reaching it is unusual.

One consequence is worth stating outright, because replacing a hardcoded threshold with this comparison does not on its own fix anything: nothing exceeds the ceiling until a confirmed spam report exists for that text. An integration that never reports a correction back has an > max_confidence branch that is just as unreachable as the >= 0.85 one it replaced. POST /v1/feedback is what makes that band reachable.

Read max_confidence from the response rather than copying the number. It is deliberately temporary: it rises when a better checkpoint lands, and a copy in your code goes stale without telling you — silently, and looking like a model change rather than a constant change. It is also per-language, so one hardcoded value is already the wrong answer for some of your traffic the moment the ceilings diverge. The response carries the ceiling that was actually applied to the text you sent.

Where to next