privacy
Privacy Policy
Last updated: September 27, 2026
Siftfy classifies the text you send to our API and returns a spam score. Standalone prediction text is normally processed in memory. When the paid shadow trial is enabled, selected paid submissions are retained for 30 days and sent to TypeSafe for a background comparison, and, when live escalation is enabled, paid submissions our model is unsure about are sent to TypeSafe on the request path and TypeSafe decides them, as described below. When evaluation collection is enabled, Content Gate submissions are retained in full for product evaluation for 30 days, with credentials removed before storage. This changed on September 19, 2026: we previously masked identifiers such as email addresses and links, and the resulting text could not be reviewed for accuracy. Records collected before September 19, 2026 remain masked. We do not write submission text to our application logs and we do not sell it. From 2026-09-27, Content Gate submissions from Siftfy's own sites that our reviewer has labeled may be used to train our models, as described under Training on labeled submissions. Separately, on September 21, 2026 we published the terms on which one narrow set of them could be used for training — submissions a customer's own honeypot field caught and reported to us as spam — if that customer authorized it in writing. No customer has, and today none can: every site running Content Gate is operated by VectraSEO LLC, the company that operates Siftfy, so there is no second party to give the authorization. That set is described under Spam label corpus, which is disclosed in full and holds no records. This policy explains exactly what we do and do not keep, where we process it, and how long we hold it.
What we collect
- Account email. Used to send magic-link sign-in tokens and account notices.
- API request metadata. Timestamp, request ID, response status, duration, detected language, the resulting spam probability, your tenant ID, and API key ID. We use it to bill, rate-limit, debug, and measure quality. This metadata does not include the text you submitted.
- Billing details. If you upgrade to a paid plan, payment information is collected and stored by our processor (Stripe). Siftfy receives only a customer reference, plan, and renewal status — never full card numbers.
- Server logs. Standard request logs including IP address, user agent, and the path requested. Retained for security and abuse prevention.
- Content Gate visitor metadata. If a customer enables Content Gate on its form, the visitor's browser sends us the registered site key, exact origin, action, a one-way content digest, proof timing, IP address, and user agent. The customer's backend then sends the submitted text for classification and, when collection is enabled, that submission text is retained in full for evaluation across customer sites. Customers are responsible for giving their visitors any notice required by law, and should account for full-text retention when they do.
- Content Gate submitter identifier. Since September 22, 2026, to let a customer refuse a visitor who submits strangers' addresses, we derive a pseudonymous identifier from the visitor's IP address (the full IPv4 address, or the IPv6 /64): a keyed hash, never the address itself, and a different identifier on every customer site. We keep it for 30 days with each verification a customer can report, and for 30 days with each report a customer files. When a customer denies a visitor, or reports one, the identifier is kept for up to 90 days on that site's deny list, until the entry expires. Reports reach other customers only as a count of how many site owners reported that identifier, never who reported it or why, and a customer's site acts on that count only if the customer turns it on. We do not use it to train any model.
- Content Gate submission text. When a customer enables evaluation collection, we retain the text a visitor submitted through that customer's form — as submitted, with credentials removed — for 30 days. This is the most sensitive category we keep; it is described in full under Content Gate evaluation collection below. From 2026-09-27, a submission from one of Siftfy's own sites that our reviewer has labeled may also be kept for 12 months to train our models, as described under Training on labeled submissions.
- Honeypot-labelled spam submissions. Separately, when a customer's own honeypot field catches a submission and the customer reports it to us as spam, we would retain that submission text for 12 months and use it to train and evaluate our spam model, together with a smaller derived record that holds no submission text and has no fixed end date. This collection is disclosed but not running: it requires that customer's separate written authorization, no customer has given one, and the corpus holds no records. It is a different collection, for a different purpose, with a different retention period, and it is described under Spam label corpus below.
- Cookies. A session cookie after sign-in (HTTP-only, 7-day expiry) and, only if you opt in, first-party Google Analytics cookies on our marketing website. We save your analytics preference in your browser and may use an essential preference cookie to keep analytics off if browser storage is unavailable. You can reject analytics or change your choice at any time using “Analytics preferences” in the footer. We use no advertising or cross-site tracking cookies.
What we do not store
- Prediction text outside the paid shadow trial. Text sent to
/v1/predictis processed in memory and discarded unless selected for the enabled paid shadow trial described below. No submission text is included in logs, error reports, or support systems. The trial is for evaluation, not model training. - Credentials inside Content Gate text. Before evaluation storage we remove private key blocks, bearer tokens, and values labelled as a password, secret, API key, or token. We do not remove anything else: the retained text is the submission as it arrived, including any email address, phone number, link, or other personal information in it. We treat it as restricted personal data, never as anonymous data.
- Request content in diagnostics. Our logs and validation errors are built to carry only metadata (status, timing, language, probability, identifiers). They do not echo the submitted text.
- Content Gate challenge material. We do not log submission text or persist content digests, challenge tokens, proof nonces, verification secrets, authorization values, or raw idempotency keys. Evaluation storage receives the submission text with credentials removed; it receives no request headers, credentials, or challenge material. The browser-facing endpoint never receives the score or decision.
Content Gate evaluation collection
When enabled, we collect the submission text from first-time authenticated Content Gate verifications across customer sites, including allowed, reviewed, blocked, and indeterminate outcomes. A submission whose entire text is a single email address, such as a sign-in link request, has nothing to evaluate and is not collected. Alongside it we retain keyed pseudonyms for customer, site, submission and text group; the form action the submission came from; the language the classifier detected; collection and expiry times; the content-handling version; and the available model score, model version and allow/review/block outcome. Authorized reviewers may add legitimate, spam, or uncertain labels to measure quality. We do not retain IP addresses, origins, authentication credentials, or challenge material in this dataset.
Evaluation records are stored in a separate private encrypted S3 bucket in the United States. Access is restricted to authorized operators and reviewers. Each record carries an expiry exactly 30 days after collection; readers exclude it from that moment, and the storage lifecycle deletes the object. Physical deletion is asynchronous and typically completes within a few days of expiry. Labels linked to expired records are excluded from evaluation. From 2026-09-27, labeled records from Siftfy's own sites may be copied into the separate training store described under Training on labeled submissions, where they are kept for 12 months; nothing in this dataset is copied into the spam label corpus — see Spam label corpus, which is a separate collection that holds no records.
The same bucket also holds records of collection failures. When one of our servers shuts down having failed to store submissions it had accepted, it writes down how many it did not store and its own server name, so that a gap in this dataset is visible to us rather than assumed not to exist. These reports contain no submission text and no customer, site or visitor identifier, and they are deleted 31 days after they are written.
Deletion on request works at the level of a customer and a site, which is the only level at which we can find these records: customer, site and submission identities are stored as one-way keyed pseudonyms and the text is not indexed by any person's identity. We cannot locate one visitor's submission on our own. A visitor who wants their submission deleted should ask the customer whose form they used; that customer can ask us to erase its records, and we will. Email hi@siftfy.io.
Why we changed this: masked text could not be read well enough to tell a correct decision from an incorrect one, so we could not measure or improve accuracy from it. Keeping the submission raises the consequence of a breach of this store, and the 30-day window, the restricted access, and the removal of credentials are what we do about that.
Paid TypeSafe shadow evaluation
This optional operator-controlled trial compares new Pro submissions from /v1/predict and Content Gate in the background, for evaluation. The comparison is recorded after the response; the live escalation described below is the path on which TypeSafe decides. Free-tier traffic and historical records without verified paid status are excluded. The initial selection uses a local score from 0.15 to 0.85 and no more than 20,000 UTF-8 bytes of text. The score band is an evaluation rule, not a calibrated confidence estimate.
When enabled, selected prediction text is retained in the same private evaluation storage for 30 days from collection, with credentials removed, account and submission identifiers pseudonymized, and access restricted to authorized reviewers. It may still contain personal information. Reviewers may save legitimate, spam, or uncertain labels. TypeSafe AI, Inc. receives the credential-stripped text and fixed classification questions. We do not send our human labels, local scores, account identifiers, request headers or authentication credentials to TypeSafe. Trial text from /v1/predict is not added to any training store, and TypeSafe's answers are never used for training; a labeled Content Gate submission from Siftfy's own sites is handled as described under Training on labeled submissions.
TypeSafe processing is governed by its data processing terms and privacy policy. Its policy states that inputs are not used to train or fine-tune models. We do not claim zero retention or a fixed retention period at TypeSafe; Siftfy's 30-day storage limit is not a statement about the provider's retention. Provider processing locations and transfers follow those terms, rather than a Siftfy-only US storage guarantee.
Comparison results hold a submission pseudonym, model and prompt versions, disposition, confidence, token usage, timing and local score, with no second copy of the text. They become unavailable when the source expires and have a 30-day physical lifecycle from creation. New admin group labels also have a 30-day lifecycle. Daily budget counts and operator control records contain no submission text and persist to prevent resets of the allowance. Customers may request erasure of their evaluation records, labels and comparisons through hi@siftfy.io.
Live TypeSafe decisions for paid submissions
When live escalation is enabled, TypeSafe decides paid submissions our own model is unsure about, and its answer changes what we return. It is enabled first only for Content Gate sites and API keys that Siftfy's operator runs itself. It is not enabled for another customer's account until that customer has had at least 30 days' notice. Free-tier traffic is never sent.
A paid Content Gate submission whose local score is at or above the site's review threshold, up to and including the model's published confidence ceiling, and a paid /v1/predict request whose local score is from 0.5 up to that ceiling, is sent to TypeSafe AI, Inc. with credentials removed, together with fixed classification questions, while the request waits for up to one second. When TypeSafe gives a confident spam or legitimate answer, that answer is the Content Gate decision (block or allow, with reason: "typesafe") or replaces spam_probability on /v1/predict (with score_source: "typesafe"). Otherwise we return what our own model alone would have returned, and the local score is unchanged. A submission whose whole text is one email address is never sent. We do not send human labels, local scores, account identifiers, request headers or authentication credentials.
Each live answer is recorded like a shadow comparison, with no second copy of the text, and the same 30-day limits apply. Where evaluation collection is enabled, the submission text of a live-escalated Content Gate submission is kept as described above, and the text of a live-escalated paid /v1/predict request is kept in the same private evaluation storage for 30 days from collection, with credentials removed, so reviewers can label it. TypeSafe's processing follows the provider terms described under Paid TypeSafe shadow evaluation. Live spend is capped per day; once the cap is reached, no further text is sent that day.
Training on labeled submissions from Siftfy's own sites
From 2026-09-27, Content Gate submissions that our reviewer has labeled may be used to train Siftfy's own models: the spam model, the AI-authorship model and the automation (bot-detection) model. This applies only to sites that VectraSEO LLC, the company that operates Siftfy, runs itself. Today that is every site running Content Gate, including Emcognito and Folio. Another customer's submissions are not used for this; training on a customer's content still requires that customer's separate written authorization, as the Terms say.
Only submissions a reviewer has labeled legitimate or spam are used. Unlabeled submissions and submissions labeled uncertain are not. A labeled submission is copied, with its label, into a separate training store in the same private encrypted storage in the United States, and kept for 12 months from collection; the storage lifecycle then deletes it. The copy is made by a daily job, only while the submission's 30-day evaluation record still exists. Models are trained from that copy only on private, self-hosted build machines that VectraSEO LLC operates in the United States and uses for nothing but its own projects. The copy has credentials removed by the same rules described above and nothing else removed, so it is personal information: it is not anonymized, not deidentified and not aggregated. Text from /v1/predict is not copied into it.
One exception reaches back before 2026-09-27. Submissions our reviewer labeled between 2026-09-22 and 2026-09-27 are also used for training, although they were collected while this page said we did not train on the evaluation collection. That departs from the principle we set out for the spam label corpus, and we are saying so here rather than leaving it to be inferred. It covers only submissions from Siftfy's own sites that a reviewer labeled in those days; no unlabeled submission collected before 2026-09-27 is used.
TypeSafe output is not used for training. TypeSafe's terms do not permit it. Our reviewer sees TypeSafe's verdict on a submission only after saving a label, and the training copy stores no TypeSafe answer.
Spam label corpus
This corpus is disclosed and specified, and it is not collecting. It holds no records. There is no path in the running service that puts a record into it, and no customer has authorized one. This section sets out the terms it would run on if that ever changed, so that those terms are published and dated before any collection rather than after it.
Why no customer has authorized it, and why today none can. Training here is conditional on a customer's separate written authorization. Every site running Content Gate is operated by VectraSEO LLC, the company that operates Siftfy. An authorization is an instruction one party gives another, and we cannot give one to ourselves, so the set of parties able to start this collection is currently empty. It would become available if a customer unaffiliated with us enabled Content Gate and authorized it in writing. Earlier on September 21, 2026 this page said we had begun training on this set. We had not begun, and no record was ever collected. That statement was wrong when we published it; this paragraph is the correction, not a change of practice.
Nothing collected before September 21, 2026 is in this corpus, and nothing collected before that date will be added to it. Those submissions reached us while we were saying we would not train on them, and a later version of a privacy policy is not a licence to go back and use data on terms nobody was given at the time. The cut-off is a constant in our code rather than a sentence in this policy: a record dated before it is refused by the software, not by our intentions.
What goes in it
Only submissions a customer's own honeypot caught. A honeypot is a form field a human visitor never sees and never fills in, so anything that fills it is automated. When a customer's form traps a submission that way and the customer reports it to us as spam, that submission becomes a label. We do not add submissions that only our own model thought were spam, we do not add anything a person at Siftfy guessed at, and we do not add a customer's submissions at all unless that customer has separately authorized it in writing.
What we keep, for how long, and why that differs
- The submission text — 12 months. Stored as it arrived, with credential material removed by the same rules described above and nothing else removed. It contains whatever the sender put into the form, which routinely includes email addresses, phone numbers, links and names. We would use it to train and evaluate our spam model. It would be personal information. It is not anonymized, not deidentified, not aggregated, and we will not describe it as any of those things: identifiers are still in the text, and one-way pseudonyms on the customer and site fields do not make the text itself anonymous. The 12-month window is a constant in our code, so a record that would sit here for longer cannot be written at all.
- A derived record with no submission text — no fixed end date. When the text expires at 12 months we keep a small record of the fact that the label existed. It holds these fields and no others: the verdict; the label source; the name of the honeypot field that was filled, never its value; the form timer delta; the customer's own disposition of the submission; when it was observed; the customer's reference for it, stored as a one-way keyed pseudonym rather than as given; and the spam probability and model version our processor produced at the time. There is no field for submission text. That is a property of the record's structure rather than a promise about our conduct — the type has nowhere to put any, and a test fails the build if someone adds one.
Why the derived record has no end date, since we have to tell you the criteria if we cannot give you a period. We keep it for as long as the model whose behaviour it documents, or a model trained from it, is in service, plus the period in which we may have to answer for decisions that model made. It is the only durable evidence of how a deployed model scored a message a human had already judged to be spam, and that question stops being answerable if the evidence expires with the text. When a model line is retired and no longer subject to inquiry, its derived records go with it. To be exact about the mechanism rather than implying a timer: the 12-month text is deleted automatically by a storage rule, whereas deleting a derived record is an action a person takes against the criteria above, because whether a model is still in service is not something storage can work out on its own. This record is pseudonymized, not anonymous: the identifiers in it are one-way keyed values, and we do not claim they could never be linked back to anyone.
Where it goes, and where it does not
The corpus is provisioned in the same private encrypted storage in the United States as the evaluation collection, under separate prefixes and separate lifecycle rules so that neither window can delete the other's records early. The prefix exists, its twelve-month rule is live, and it is empty — the window is in place so that the published figure is true from the first record rather than configured after one arrives. Access is restricted to authorized operators. No submission text from this corpus is sent to any third-party model provider. We do use an external model API to pre-label a separate training corpus of public and self-owned content; records from this corpus are excluded from it, and that exclusion is a pre-flight check that refuses to send anything at all if a single forbidden record is present. Nothing here is sold, shared for advertising, or disclosed to another customer.
Deletion
As with the evaluation collection, we can find these records at the level of a customer and a site, and not at the level of one person — the text is not indexed by anybody's identity, so we cannot locate a single sender's submission on our own. If you submitted a form and want what you sent removed, ask the operator of the site you submitted it to; they can require us to erase their records, including their labels and the derived records behind them, and we will. Customers can write to hi@siftfy.io. A customer that had given the authorization and then withdrew it would stop contributing new labels immediately; already-trained model weights cannot be unlearned record by record, and we will not pretend otherwise.
Feedback corrections
If you use the optional /v1/feedback endpoint to tell us a message was correctly or incorrectly classified ("ham" or "spam"), we use the text you send only to compute a salted one-way fingerprint of it. We store that fingerprint, the ham/spam label, your tenant ID, and a count of how many times the same correction was submitted — never the readable text itself. The fingerprint cannot be reversed back into your content; it exists only so we can recognise the same correction again and improve classification accuracy. This is a correction signal tied to your account, and it is deleted with your account (see Your rights). Outside the enabled paid shadow trial, standalone prediction traffic is not retained. Neither standalone prediction content nor feedback text is used for model training. The one collection whose terms would permit it is the Spam label corpus described below, and that corpus is not collecting.
How we use what we collect
- Authenticate you and protect your account.
- Bill, rate-limit, and operate the API.
- Detect abuse, debug failures, and investigate security incidents.
- Measure classifier quality in aggregate (for example, score distribution by language) using metadata only.
- Send transactional email tied to your account (sign-in, billing receipts, plan changes).
Where we process data
Siftfy-operated application compute and storage are in the United States. TypeSafe processing during the paid shadow trial and live escalation is described above and follows the provider terms. Our databases, transactional email, and storage run on Amazon Web Services in the US (region us-east-1, Northern Virginia), and our application compute runs in a US region. If you access Siftfy from outside the United States, you are sending your data to the United States for processing.
Sharing and subprocessors
We share data with infrastructure providers only as necessary to operate the service. We do not sell or rent your data, and we do not share it with advertisers. Our current subprocessors are:
- Amazon Web Services — hosting, database (DynamoDB), and transactional email (SES); United States.
- TypeSafe AI, Inc. — classification of selected paid submission text: in the background during the enabled shadow trial, and as the live decision on paid submissions our model is unsure about when live escalation is enabled; provider locations, transfers and retention follow its data processing terms.
- Stripe — payment processing and subscription billing.
- Google Analytics — first-party, opt-in usage analytics on the marketing website only; never receives API request content.
- Fontshare — web-font delivery on the marketing website.
The canonical, current list lives on our Trust & Security page. We will disclose data to authorities only if compelled by valid legal process and, where permitted, will notify the affected account.
Retention
- Content Gate evaluation text: 30 days from collection, then hard-deleted; physical S3 removal follows asynchronously. Records collected before September 19, 2026 hold masked text and keep their original window of up to 90 days. Imported labels are retained for up to 90 days from import and cease contributing when their source expires.
- Spam label corpus — submission text: 12 months from collection, then hard-deleted; physical S3 removal follows asynchronously. Nothing collected before September 21, 2026 is eligible. The corpus holds no records, so this is the window that would apply rather than one now running. See Spam label corpus.
- Spam label corpus — derived record (no submission text): no fixed period. Retained while the model it documents, or a model trained from it, is in service and answerable for its decisions; deleted when that model line is retired. The fields are enumerated exhaustively under Spam label corpus.
- Content Gate collection-failure reports: 31 days from the report, then deleted. These hold a count and a server name only — no submission text and no identifiers.
- Standalone prediction request content: processed in memory and discarded outside the enabled paid shadow trial and live escalation. Selected trial text, and live-escalated paid text where evaluation collection is enabled, is retained by Siftfy for 30 days from collection, with credentials removed; provider processing is described under Paid TypeSafe shadow evaluation and Live TypeSafe decisions.
- Content Gate verification records: a content-free terminal result and one-way replay/idempotency hashes are logically expired after approximately 10 minutes; DynamoDB may physically remove expired rows later. Site configuration remains until the customer deletes the site or account.
- Account email and account metadata: retained for the life of the account.
- Usage counts: retained while the account is active to support billing.
- Server and application logs: retained for up to 90 days.
- Revoked API keys: retained for 90 days after revocation, then purged.
- Magic-link sign-in tokens: expire within 15 minutes.
- Feedback fingerprints: retained as accumulated correction signal; contain no readable content and are deleted with your account.
- Billing records: retained as required by tax and accounting law (typically 7 years).
- Backups: point-in-time and versioned backups roll off within approximately 30 days.
Your rights
You may request access to, a copy of, correction of, or deletion of your account data. Email hi@siftfy.io from the address on the account and we will respond within 30 days. On a deletion request we remove your account email, API keys, passkeys, Content Gate sites, usage records, and the feedback fingerprints attributed to your tenant; backups containing that data roll off within approximately 30 days. Billing records we are legally required to keep are retained for the applicable period. Depending on where you live, you may have additional rights under laws such as the GDPR/UK GDPR or the CCPA/CPRA; we honour valid requests under those laws and do not discriminate against you for exercising them.
International transfers and DPA
For the limited request metadata described above, Siftfy acts as a data processor on your behalf, and our data-processing terms are set out in this policy. If your organization requires a separate Data Processing Agreement, contact hi@siftfy.io and we will work with you to put appropriate terms in place, including the Standard Contractual Clauses where they are required for transfers from the EEA, the UK, or Switzerland to the United States. See the Trust & Security page.
Security
Traffic is TLS-only. Passwords are not stored — sign-in is via magic-link email or WebAuthn passkey. API keys are stored hashed and shown to you only once at creation. We follow the principle of least privilege for internal access and review production access on a recurring basis. More detail is on our Trust & Security page.
Children
Siftfy is a developer tool and is not directed at children under 13, and we do not knowingly collect their data.
Changes
We will update this page with material changes and refresh the "last updated" date above. We will not apply a materially more permissive use to data we already collected — for example, beginning to train on previously submitted content — without first giving notice and, where required, obtaining your consent. The change dated 2026-09-27 is one such change, and it is limited to Siftfy's own sites: submissions our reviewer labeled from 2026-09-22 are used for training, as described under Training on labeled submissions. This page and the privacy notices of those sites are the notice, and none of those submissions is copied for training before they are published. VectraSEO LLC is the customer for those sites; we have not asked the people who filled in their forms for consent. Continued use of the service after a change constitutes acceptance of the revised policy.
Contact
Siftfy is operated by VectraSEO LLC, a Pennsylvania limited liability company, which is the controller for the account and billing data described above. Questions about this policy? Email hi@siftfy.io.