Training & Data Use
Last updated: 2026-03-15
ScanLens may use anonymized scan events to improve detection quality and to train internal machine‑learning models. The goal is to learn general patterns of malicious links (phishing / scams / malware) without storing sensitive user data.
What is stored for training
- Hashed URL identifiers (non‑reversible hashes) so we can deduplicate and measure drift.
- Derived URL features such as length, number of dots/hyphens/digits, presence of punycode, and whether the host looks like an IP address.
- Outcome labels such as verdict (SAFE / SUSPICIOUS / DANGEROUS), risk score, and reason codes/signals.
- Timing (timestamps) to analyze trends and abuse waves.
What is NOT stored for training
- Full raw URLs (including paths and query parameters).
- Query parameter values (tokens, emails, IDs, etc.).
- User‑provided text content from webpages.
- Personal identifiers like name, email, phone number (there is no account system).
Why hashing & derived features
We intentionally prefer non‑reversible representations. Hashes let us detect repeats and measure prevalence. Derived features let models learn statistical patterns while avoiding storage of sensitive strings.
Retention
Training logs are rotated and retained for a limited period (for example ~30 days) before being deleted. Aggregated and de‑identified metrics may be retained longer to measure performance and improve the product.
Opt‑out
If you need an opt‑out for training use, use the Support page or email [email protected]. Note that we still must process the URL to return a scan result.