benchmarks · methodology

Content spam detection benchmark notes.

Updated September 16, 2026

In a local run on 1,741 deduplicated UCI YouTube comments, Siftfy achieved 87.4% F1, 96.3% precision, and 80.0% recall at the 0.50 review threshold. It sent 26 of 904 legitimate comments to review (2.9%) and allowed 167 of 837 spam comments. No content was blocked by the model alone. Overlap with prior calibration data is unverified; these results do not establish accuracy on unseen customer traffic.

F1

87.4%

Local evaluation of 1,741 deduplicated YouTube comments.

Precision

96.3%

At a 0.50 review threshold; review is not blocking.

Recall

80.0%

670 of 837 labeled spam comments sent to review.

Recommended shadow run

7 days

Validate thresholds against your own traffic before enforcing.

Review band

0.50-0.84

Queue suspicious submissions instead of auto-dropping them.

What the benchmark measures

This dataset contains comments from five music videos. It does not establish accuracy on contact forms, signup notes, AI-generated spam, or supported languages. The local run exercised challenge issuance, proof solving, and verification in both monitor and enforce modes, using the default cal3 model with an empty feedback store.

F1 balances precision and recall against labeled ham and spam examples. Production teams should still inspect both metrics, especially false positives, at the thresholds they plan to enforce.

Source: Alberto and Lochter (2015), UCI YouTube Spam Collection, licensed under CC BY 4.0. We removed duplicate comments and source author/date metadata, preserving message text and labels.

How to validate it on your traffic

  1. Run Siftfy in shadow for at least one normal traffic week.
  2. Log the returned probability, your current moderation outcome, and a human review label where possible.
  3. Choose a review threshold on validation data, then measure it on a separate held-out set. Preserve the model-only blocking restriction.
  4. Manually inspect false positives before blocking anything automatically.
  5. Recheck thresholds after a spam wave, launch, or major content-type change.

Recommended threshold bands

ProbabilitySuggested actionWhy
0.85-1.00Block only with trusted evidenceModel and lexical evidence alone cannot reach this band.
0.50-0.84Review queueBorderline text where false positives are more likely.
0.00-0.49Allow, with normal abuse monitoringMost legitimate submissions should live here.

Known limits

  • Legitimate links can trigger review; some scams and multilingual spam can be missed.
  • Very short messages have less signal and may need stricter product-side rules.
  • The API scores content, not IP reputation, disposable email status, device fingerprint, or payment risk.
  • Spam tactics change, so thresholds should be reviewed regularly instead of frozen forever.

To test the API, read the predict endpoint docs, try the spam probability tester, or compare providers in the spam detection API guide.

Download the aggregate evaluation results and model manifest, and see how to keep latency low in optimizing spam-detection API performance.