"Found 3,000 duplicate photos in under 5 minutes. No install, no upload — my files never left my laptop."
How It Works
- 1. Pick a FolderSelect any folder on your computer — subfolders are included automatically
- 2. Scan RunsSakarto processes every file and builds fingerprints inside your browser. No data leaves your machine.
- 3. Groups Appear LiveDuplicate groups appear as they are found — no waiting for the full scan to finish
- 4. Move or DeleteSelect files from any group and move, delete, or compare them side-by-side
Image & Video Duplicate Finder — 7 Algorithms
Scan a folder and Sakarto groups all visually similar images and videos. Every algorithm produces a 64-bit hash compared using Hamming distance — lower threshold = stricter matching. Pick based on what kind of duplicates you expect.
| Algorithm | How it works | Canvas / Basis | Rotation-safe | Re-compression | Brightness shifts | Speed |
|---|---|---|---|---|---|---|
| Color Signature | Compares YUV colour values across a 24×24 regional grid. The only algorithm that actually compares colour, not just luminance. | 24×24 YUV grid | — | ✓ Good | ⚠ Partial | ⚡ Fast |
| aHash | Resizes to 16×16, converts to grayscale, compares each pixel to the overall mean. The simplest and fastest hash. | 16×16 grayscale | — | ✓ Good | ⚠ Moderate | ⚡ Fastest |
| BlockHash | Divides a 64×64 canvas into 16×16 blocks, averages each block, compares to median. Block averaging absorbs JPEG noise and compression artifacts. | 64×64, 16 blocks | — | ✓ Excellent | ✓ Good | ⚡ Very Fast |
| dHash | Resizes to 17×16 (one extra column), compares each pixel to the one on its right. Encodes horizontal brightness gradients, not absolute values. | 17×16 gradient | — | ✓ Good | ✓ Excellent | ⚡ Very Fast |
| pHash | Resizes to 32×32, applies a 2D Discrete Cosine Transform, extracts top-left 8×8 low-frequency coefficients, thresholds against mean. Frequency-domain encoding is very stable across edits. | 32×32 DCT | — | ✓ Excellent | ✓ Good | ⚠ Moderate |
| wHash | Resizes to 16×16, applies a Haar Wavelet Transform row-wise then column-wise, extracts 8×8 LL sub-band, thresholds against mean. Similar quality to pHash but faster. | 16×16 Haar | — | ✓ Good | ✓ Good | ⚡ Fast |
| ORB | Uses OpenCV.js WASM to detect FAST keypoints and compute BRIEF descriptors. Matches keypoints spatially using Hamming distance + Lowe's ratio test. Threshold is inverted: higher = stricter. | Keypoint features | ✓ | ✓ Good | ✓ Good | Slower — O(n²) |
Color Signature Duplicate Finder
The only algorithm that compares actual colour — YUV values across a 24×24 regional grid. Finds copies sharing the same colour palette even when re-encoded or slightly edited. Colour-shifted versions of the same image will not match.
Average Hash (aHash) Duplicate Finder
The fastest duplicate finder. Resizes to 16×16, converts to grayscale, and compares each pixel to the overall mean — groups files sharing the same broad brightness distribution. Best for exact duplicates and resized copies.
Block Hash Duplicate Finder
Divides a 64×64 canvas into 16×16 blocks and averages each. Block averaging absorbs JPEG compression noise and image artifacts, making it more tolerant of degraded copies than pixel-level hash methods.
Difference Hash (dHash) Duplicate Finder
Resizes to 17×16 (one extra column for gradients), converts to grayscale, and compares each pixel to the one on its right. Encodes relative brightness direction — not absolute values — making it robust to brightness and exposure changes.
Perceptual Hash (pHash) Duplicate Finder
Resizes to 32×32, applies a separable 2D Discrete Cosine Transform, and extracts the top-left 8×8 low-frequency coefficients. Frequency-domain encoding is very stable across re-compression, mild colour grading, sharpening, and format conversion.
Wavelet Hash (wHash) Duplicate Finder
Resizes to 16×16, applies a Haar Wavelet Transform (averages and differences row-wise then column-wise), and extracts the 8×8 LL sub-band. Produces quality similar to pHash at lower CPU cost — the best speed/quality balance.
Feature Matching (ORB) Duplicate Finder
Uses OpenCV.js (WebAssembly) to detect FAST keypoints and compute BRIEF binary descriptors. Matches descriptors with Hamming distance + Lowe's ratio test. The only algorithm that handles rotation, perspective distortion, and cropping. Supports images and videos — 3 frames are extracted per video and the frame with the most keypoints is used.
📊 Which Image / Video Algorithm?
Quick reference for choosing based on what kind of duplicates you expect to find.
| I want to find… | Color Sig. | aHash | BlockHash | dHash | pHash | wHash | ORB |
|---|---|---|---|---|---|---|---|
| Exact byte-for-byte copies | ✓ | Best | ✓ | ✓ | ✓ | ✓ | ✓ |
| Resized versions (different resolution) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Re-compressed / different format | ✓ | ✓ | Best | ✓ | ✓ | ✓ | ✓ |
| Brightness / exposure adjusted | — | ⚠ | ⚠ | Best | ✓ | ✓ | ✓ |
| Colour-graded (hue/saturation changed) | — | ⚠ | ⚠ | ✓ | ✓ | ✓ | ✓ |
| Same colour palette / scene | Best | — | — | — | — | — | — |
| Watermarked or lightly edited | ⚠ | ⚠ | ⚠ | ✓ | Best | ✓ | ✓ |
| Rotated or mirrored copies | — | — | — | — | — | — | Best |
| Cropped or perspective-warped | — | — | — | — | — | — | Best |
| Very large folder (speed priority) | — | Best | ✓ | ✓ | ⚠ | ✓ | Avoid |
| Videos as well as images | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ⚠ Best-keypoint frame |
Audio Duplicate Finder — 3 Algorithms
Scan a folder and Sakarto groups all duplicate and similar audio files. Three algorithms, each measuring a different quality of sound. All audio processing continues in the background even when you switch browser tabs.
| Algorithm | What it measures | Best for | Finds covers / alt versions | Finds speech / podcasts | Speed | External library |
|---|---|---|---|---|---|---|
| Chromaprint Spectral hash | Spectral band energy hashed into compact 15-bit frames. Captures overall frequency shape of the audio. Pure JavaScript, no CDN dependency. Lower threshold = stricter. | Exact copies, re-encoded files, different bitrates of the same master track | ⚠ Limited | ✓ Good | ⚡ Fastest — pure JS | ✓ None |
| Essentia HPCP Harmonic chroma | 12-bin Harmonic Pitch Class Profile per frame via Essentia.js WASM at 44,100 Hz. Captures chord and pitch class content. Compares using cosine similarity. Higher threshold = stricter. | Music with harmonic similarity — covers, transpositions, live vs studio recordings | ✓ Excellent | ✗ Poor | ⚠ Moderate — WASM init 1–3s once | ⚠ Essentia.js (CDN) |
| Meyda MFCC Timbral texture | 13 Mel-Frequency Cepstral Coefficients per frame via Meyda.js at 22,050 Hz. Analyses first 10 seconds only. Captures the perceptual "colour" of sound. Lower threshold = stricter. | Podcasts, voice memos, sound effects — audio where timbre matters more than pitch content | ✗ Poor | ✓ Excellent | ⚡ Fast — 10s window per file | ⚠ Meyda.js (CDN) |
Chromaprint Audio Duplicate Finder
Pure JavaScript spectral fingerprinting. Hashes spectral band energy into compact 15-bit frames and groups files using Hamming distance. No external library or CDN needed — works offline. Analyses the full audio track.
Essentia HPCP Audio Duplicate Finder
Uses Essentia.js WASM to extract Harmonic Pitch Class Profiles — 12-bin chroma vectors at 44,100 Hz. Groups files by cosine similarity of their harmonic content. Loads the WASM library once per session. Threshold direction: higher = stricter.
Meyda MFCC Audio Duplicate Finder
Uses Meyda.js to extract 13 Mel-Frequency Cepstral Coefficients at 22,050 Hz. Analyses only the first 10 seconds of each file — very fast. Groups files by Euclidean distance of their timbral texture. Captures how sound "feels" to the ear, independent of pitch.
⭐ What Users Say
"The Chromaprint audio finder is incredible. Cleaned up my entire podcast archive in one session."
"The ORB algorithm found rotated scans I had no idea were duplicates. Nothing else catches those."
"pHash found hundreds of re-compressed duplicates I'd been accumulating for years. Privacy is a huge plus too."