28 Commits

Author SHA1 Message Date
Sreedev Kodichath
2a5324b48c [ADD] tui: interactive terminal UI for resolving duplicates
The tool could only act on duplicates non-interactively or through a
line-oriented per-group prompt. Add a full-screen terminal UI (--tui) for
browsing and resolving them visually.

The UI is a two-pane master/detail view: a list of duplicate groups and,
for the selected group, its files with per-file deletion marks. Files can
be marked individually or in bulk by a keep-strategy (newest/oldest/first/
last/shortest/shallowest, reused from the resolver) applied to the current
group or to all groups at once. A footer tracks how many files are marked
and how much space would be reclaimed. Deletion is gated behind an explicit
confirmation, a group always keeps at least one file, and the highlighted
file can be opened in the system default application.

The design keeps the interaction logic pure and independently testable: an
App model translates an abstract key event into an outcome, so navigation,
marking, strategy application, and post-deletion model updates are all unit
tested without a terminal. Rendering (ratatui) and the destructive file I/O
live in the event loop, which restores the terminal on both normal exit and
panic so a crash never leaves it broken. --tui is mutually exclusive with
--keep and --interactive.
2026-07-26 22:02:17 +02:00
Sreedev Kodichath
1442cd0aaf [ADD] cache: on-disk hash cache for faster re-runs
Content hashing is the pipeline's expensive stage; every run re-read and
re-hashed all size-collision candidates from scratch. Persist those hashes
so an unchanged file is skipped on subsequent runs.

Add a path-keyed cache mapping absolute_path -> (size, mtime, strict, hash,
last_seen), stored in a compact hand-rolled little-endian binary format with
a magic+version header. A candidate reuses its stored hash only when size,
mtime, and hash-mode all match; otherwise it is re-hashed and the entry is
refreshed. On save, entries unseen for 30 days are pruned so the file cannot
grow without bound. The hashing stage splits into hash_candidates (parallel,
cache-consulting) and group_hashed (grouping) so lookups stay inside the
parallel pass while cache mutation stays single-threaded after it.

The cache is strictly a performance hint: a missing, corrupt, version-
mismatched, or unwritable cache degrades to a correct full-hash run and
never fails or changes the result. Caching is on by default at the platform
cache directory; --no-cache disables it and --cache-file overrides the path.

The gxhash seed becomes a fixed constant so stored hashes are reproducible
across runs, which is what makes cache hits possible; rand is consequently
no longer a production dependency (now dev-only), and dirs is added for the
platform cache directory.
2026-07-26 20:53:41 +02:00
Sreedev Kodichath
cc7144afbc [ADD] resolver: bulk duplicate resolution via --keep/--force
The tool could only act on duplicates through the default listing or the
manual per-group interactive prompt, with no way to resolve groups in bulk
by a rule. This blocked any scripted or automated use, and bulk resolution
is the top item in the README's proposed operations.

Add a resolver module that reduces each duplicate group to a single kept
file. --keep <newest|oldest|first|last|shortest|shallowest> selects the
keeper: newest/oldest by mtime, first/last and shortest/shallowest by
path, each resolved to a unique keeper through total-order path tiebreaks
so the outcome is deterministic and order-independent. Selection is a pure,
independently tested function; deletion is a separate step.

Deletion is guarded: --keep alone previews (KEEP/DELETE lines and a
would-free summary) and removes nothing, so a mistaken invocation is
harmless. --force performs the deletion, continues past individual
failures, reports freed space, and exits non-zero if any file could not be
removed. clap enforces that --keep excludes --interactive and that --force
requires --keep, so misuse fails before any file is touched.

Also fixes the interactive table's column width, which sized on path
component count instead of path string length so the columns never aligned.
2026-07-26 19:27:49 +02:00
Sreedev Kodichath
3ca931fa8e [REF] core: replace streaming pipeline with staged batch pipeline
The deduplication core was a three-stage streaming producer/consumer
(scan -> group-by-size -> group-by-hash) orchestrated by a Server struct
that ran the stages concurrently on a threadpool. Coordination relied on
AtomicBool flags polled in busy-wait loops, an Arc<Mutex<Vec>> hand-off
queue, DashMap stores, and a per-file Arc<Mutex<FileState>>. Every stage
had to receive its own hand-cloned Arcs with ad-hoc names, and the
busy-wait branches burned a core spinning while waiting for the producer.

Replace it with a staged batch pipeline over owned collections:
pipeline::run(&Params) drives scan -> group_by_size -> group_by_hash in
sequence, with rayon supplying parallelism. With no shared mutable state
between stages, all the Arc plumbing, the AtomicBool coordination, the
Mutex queue, the per-file lock, server.rs, and the threadpool/dashmap
dependencies are gone. FileInfo becomes plain data and each stage is a
pure, independently testable function.

Behavior is preserved except for three authorized deviations: progress
spinners render sequentially rather than concurrently under -p; the
interactive empty-result message prints once instead of twice; and the
interactive "Duplicate Set X of Y" total now counts the confirmed
duplicate groups shown instead of the internal candidate-hash store size.

This lands as the foundation for the cache, CLI, and TUI work that follows.
2026-07-26 14:27:02 +02:00
sreedevk
d102d92bfa cleanup clone littering in server.rs 2025-08-07 23:18:02 +00:00
sreedevk
68c2491fd0 added file store partitioning idea to readme. 2025-07-30 13:08:54 +00:00
sreedevk
da24efbb36 count graphemes instead of chars in path length check 2025-07-30 10:47:58 +00:00
sreedevk
2191763b4d added potential optimizations to readme 2025-07-28 14:22:13 +00:00
sreedevk
b84510f18e tree rendering ui improvements 2025-07-19 16:38:17 +00:00
Sreedev Kodichath
bfb89edd16 Update README.md 2025-07-19 12:32:09 -04:00
sreedevk
21f0eb3cb0 all file type related filtering issues fixed 2025-07-19 16:26:12 +00:00
sreedevk
f5d2d4e22c fix: single file hash groups printed to screen 2025-07-19 14:22:43 +00:00
sreedevk
360af25318 updated readme 2025-07-19 13:57:26 +00:00
sreedevk
d06fe93171 fix pipeline breaks 2025-07-19 13:41:52 +00:00
sreedevk
6a24ca5ecf updated github release pipeline 2025-07-19 12:15:42 +00:00
sreedevk
6f23a51a53 created a automatic release ci pipeline 2025-07-19 12:00:25 +00:00
sreedevk
4894d1b4db minor ui improvements 2025-07-19 11:43:26 +00:00
sreedevk
c0286b17cc fix: hash collision edge case 2025-07-19 11:33:52 +00:00
sreedevk
e78dfc4211 updated dependencies 2025-07-19 02:01:15 +00:00
sreedevk
66f49a9aee benchmarks updated 2025-07-19 01:35:24 +00:00
sreedevk
711eb36fb9 output formatting and printing performance improvements 2025-07-19 01:26:46 +00:00
sreedevk
130a8f99ba partial hash page count increased 2025-07-18 11:25:37 +00:00
sreedevk
d97af6b8f6 improved parallelization + updated benchmarks 2025-07-17 00:31:40 +00:00
sreedevk
90198502a3 added exclude file types.
updated documentation to include exclude file types.
2025-07-14 12:56:07 +00:00
sreedevk
e1b81b9981 page hash aggregation method improvement 2025-07-14 01:02:14 +00:00
sreedevk
48e7c62036 added planned features 2025-07-14 00:23:16 +00:00
Sreedev Kodichath
58e33cc8da v0.3 (#61)
- [x] parallelization
    - [x] (scanning) + (processing sw & processing hw & formatting & printing)
- [x] reduce cloning values on the heap
- [x] add a partial hashing mode (--strict)
- [x] add unit tests
- [x] add silent mode
- [x] update documentation
- [x] remove color output
- [x] progress bar improvements
    - [x] use progress bar groups
- [x] remove broken json rendering
- [x] add benchmarks
2025-07-13 19:03:14 -04:00
sreedevk
f9e86f8522 performance improvements 2025-07-12 00:24:25 +00:00
27 changed files with 3895 additions and 1679 deletions

View File

@@ -1,137 +1,121 @@
# CI that: name: Deduplicator Release Build CI Pipeline
#
# * checks for a Git Tag that looks like a release
# * creates a Github Release™ and fills in its text
# * builds artifacts with cargo-dist (executable-zips, installers)
# * uploads those artifacts to the Github Release™
#
# Note that the Github Release™ will be created before the artifacts,
# so there will be a few minutes where the release has no artifacts
# and then they will slowly trickle in, possibly failing. To make
# this more pleasant we mark the release as a "draft" until all
# artifacts have been successfully uploaded. This allows you to
# choose what to do with partial successes and avoids spamming
# anyone with notifications before the release is actually ready.
name: Release
permissions:
contents: write
# This task will run whenever you push a git tag that looks like a version
# like "v1", "v1.2.0", "v0.1.0-prerelease01", "my-app-v1.0.0", etc.
# The version will be roughly parsed as ({PACKAGE_NAME}-)?v{VERSION}, where
# PACKAGE_NAME must be the name of a Cargo package in your workspace, and VERSION
# must be a Cargo-style SemVer Version.
#
# If PACKAGE_NAME is specified, then we will create a Github Release™ for that
# package (erroring out if it doesn't have the given version or isn't cargo-dist-able).
#
# If PACKAGE_NAME isn't specified, then we will create a Github Release™ for all
# (cargo-dist-able) packages in the workspace with that version (this is mode is
# intended for workspaces with only one dist-able package, or with all dist-able
# packages versioned/released in lockstep).
#
# If you push multiple tags at once, separate instances of this workflow will
# spin up, creating an independent Github Release™ for each one.
#
# If there's a prerelease-style suffix to the version then the Github Release™
# will be marked as a prerelease.
on: on:
push: push:
tags: tags:
- '*-?v[0-9]+*' - 'v*.*.*'
jobs: jobs:
# Create the Github Release™ so the packages have something to be uploaded to lint_test:
create-release:
runs-on: ubuntu-latest runs-on: ubuntu-latest
outputs:
has-releases: ${{ steps.create-release.outputs.has-releases }}
env: env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} RUSTFLAGS: "-C target-feature=+aes,+sse2"
steps: steps:
- uses: actions/checkout@v3 - name: Checkout repository
uses: actions/checkout@v4
- name: Install Rust - name: Install Rust
run: rustup update 1.71.0 --no-self-update && rustup default 1.71.0 uses: dtolnay/rust-toolchain@stable
- name: Install cargo-dist with:
run: curl --proto '=https' --tlsv1.2 -LsSf https://github.com/axodotdev/cargo-dist/releases/download/v0.0.7/cargo-dist-installer.sh | sh components: clippy
- id: create-release
run: |
cargo dist plan --tag=${{ github.ref_name }} --output-format=json > dist-manifest.json
echo "dist plan ran successfully"
cat dist-manifest.json
# Create the Github Release™ based on what cargo-dist thinks it should be - name: Run cargo check
ANNOUNCEMENT_TITLE=$(jq --raw-output ".announcement_title" dist-manifest.json) run: cargo check
IS_PRERELEASE=$(jq --raw-output ".announcement_is_prerelease" dist-manifest.json)
jq --raw-output ".announcement_github_body" dist-manifest.json > new_dist_announcement.md
gh release create ${{ github.ref_name }} --draft --prerelease="$IS_PRERELEASE" --title="$ANNOUNCEMENT_TITLE" --notes-file=new_dist_announcement.md
echo "created announcement!"
# Upload the manifest to the Github Release™ - name: Run cargo clippy
gh release upload ${{ github.ref_name }} dist-manifest.json run: cargo clippy -- -D warnings
echo "uploaded manifest!"
# Disable all the upload-artifacts tasks if we have no actual releases - name: Run tests
HAS_RELEASES=$(jq --raw-output ".releases != null" dist-manifest.json) run: cargo test -- --test-threads=1
echo "has-releases=$HAS_RELEASES" >> "$GITHUB_OUTPUT"
# Build and packages all the things build:
upload-artifacts: needs: lint_test
# Let the initial task tell us to not run (currently very blunt) runs-on: ${{ matrix.os }}
needs: create-release
if: ${{ needs.create-release.outputs.has-releases == 'true' }}
strategy: strategy:
matrix: matrix:
# For these target platforms
include: include:
- os: macos-11 - target: aarch64-unknown-linux-gnu
dist-args: --artifacts=local --target=aarch64-apple-darwin --target=x86_64-apple-darwin os: ubuntu-latest
install-dist: curl --proto '=https' --tlsv1.2 -LsSf https://github.com/axodotdev/cargo-dist/releases/download/v0.0.7/cargo-dist-installer.sh | sh artifact_name: linux-aarch64
- os: ubuntu-20.04 rustflags: "-C target-feature=+aes,+neon"
dist-args: --artifacts=local --target=x86_64-unknown-linux-gnu - target: x86_64-unknown-linux-gnu
install-dist: curl --proto '=https' --tlsv1.2 -LsSf https://github.com/axodotdev/cargo-dist/releases/download/v0.0.7/cargo-dist-installer.sh | sh os: ubuntu-latest
- os: windows-2019 artifact_name: linux-amd64
dist-args: --artifacts=local --target=x86_64-pc-windows-msvc rustflags: "-C target-feature=+aes,+sse2"
install-dist: irm https://github.com/axodotdev/cargo-dist/releases/download/v0.0.7/cargo-dist-installer.ps1 | iex - target: x86_64-apple-darwin
os: macos-14
artifact_name: macos-amd64
rustflags: "-C target-feature=+aes,+sse2"
- target: aarch64-apple-darwin
os: macos-14
artifact_name: macos-arm64
rustflags: "-C target-feature=+aes,+neon"
runs-on: ${{ matrix.os }}
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
steps: steps:
- uses: actions/checkout@v3 - name: Checkout repository
uses: actions/checkout@v4
- name: Install Rust - name: Install Rust
run: rustup update 1.71.0 --no-self-update && rustup default 1.71.0 uses: dtolnay/rust-toolchain@stable
- name: Install cargo-dist
run: ${{ matrix.install-dist }} - name: Add target
- name: Run cargo-dist run: rustup target add ${{ matrix.target }}
# This logic is a bit janky because it's trying to be a polyglot between
# powershell and bash since this will run on windows, macos, and linux! - name: Install cross-compilation toolchain (Linux ARM only)
# The two platforms don't agree on how to talk about env vars but they if: matrix.target == 'aarch64-unknown-linux-gnu'
# do agree on 'cat' and '$()' so we use that to marshal values between commands.
run: | run: |
# Actually do builds and make zips and whatnot sudo apt-get update
cargo dist build --tag=${{ github.ref_name }} --output-format=json ${{ matrix.dist-args }} > dist-manifest.json sudo apt-get install -y gcc-aarch64-linux-gnu
echo "dist ran successfully"
cat dist-manifest.json
# Parse out what we just built and upload it to the Github Release™ - name: Build release binary
jq --raw-output ".artifacts[]?.path | select( . != null )" dist-manifest.json > uploads.txt
echo "uploading..."
cat uploads.txt
gh release upload ${{ github.ref_name }} $(cat uploads.txt)
echo "uploaded!"
# Mark the Github Release™ as a non-draft now that everything has succeeded!
publish-release:
# Only run after all the other tasks, but it's ok if upload-artifacts was skipped
needs: [create-release, upload-artifacts]
if: ${{ always() && needs.create-release.result == 'success' && (needs.upload-artifacts.result == 'skipped' || needs.upload-artifacts.result == 'success') }}
runs-on: ubuntu-latest
env: env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }} RUSTFLAGS: ${{ matrix.rustflags }}
steps: CARGO_TARGET_AARCH64_UNKNOWN_LINUX_GNU_LINKER: ${{ matrix.target == 'aarch64-unknown-linux-gnu' && 'aarch64-linux-gnu-gcc' || '' }}
- uses: actions/checkout@v3
- name: mark release as non-draft
run: | run: |
gh release edit ${{ github.ref_name }} --draft=false if [ "${{ matrix.target }}" = "aarch64-unknown-linux-gnu" ]; then
export CARGO_TARGET_AARCH64_UNKNOWN_LINUX_GNU_LINKER=aarch64-linux-gnu-gcc
fi
cargo build --release --target ${{ matrix.target }}
- name: Package binary
run: |
if [[ "${{ matrix.os }}" == macos* ]]; then
BINARY_NAME=$(find target/${{ matrix.target }}/release -maxdepth 1 -type f -perm +111 -print0 | xargs -0 basename | head -1)
else
BINARY_NAME=$(find target/${{ matrix.target }}/release -maxdepth 1 -type f -executable -print0 | xargs -0 basename | head -1)
fi
echo "Packaging binary: $BINARY_NAME"
mkdir -p release
cp target/${{ matrix.target }}/release/$BINARY_NAME release/$BINARY_NAME
tar -C release -czf ${{ matrix.artifact_name }}.tar.gz $BINARY_NAME
- name: Upload artifact
uses: actions/upload-artifact@v4
with:
name: ${{ matrix.artifact_name }}-binary
path: ${{ matrix.artifact_name }}.tar.gz
create_release:
needs: build
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Download all artifacts
uses: actions/download-artifact@v4
with:
path: artifacts
- name: Create GitHub Release
id: create_release
uses: softprops/action-gh-release@v1
with:
tag_name: ${{ github.ref_name }}
name: Release ${{ github.ref_name }}
body: "Automated release for version ${{ github.ref_name }}"
files: |
artifacts/*/*.tar.gz
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}

5
.gitignore vendored
View File

@@ -1,4 +1,7 @@
/target /target
/Cargo.lock /Cargo.lock
/.bacon-locations
/result-bin /result-bin
/.bacon-locations
/docs
/.claude
/.superpowers

1741
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -1,8 +1,8 @@
[package] [package]
name = "deduplicator" name = "deduplicator"
version = "0.2.2" version = "0.3.2"
edition = "2021" edition = "2021"
description = "find,filter,delete Duplicates" description = "find,filter and delete duplicate files"
repository = "https://github.com/sreedevk/deduplicator" repository = "https://github.com/sreedevk/deduplicator"
license = "MIT" license = "MIT"
authors = [ authors = [
@@ -11,41 +11,50 @@ authors = [
"Dhruva Sagar <dhruva.sagar@gmail.com>", "Dhruva Sagar <dhruva.sagar@gmail.com>",
] ]
# See more keys and their definitions at https://doc.rust-lang.org/cargo/reference/manifest.html [[bin]]
name = "deduplicator"
path = "src/main.rs"
# See more keys and their definitions at https://doc.rust-lang.org/cargo/reference/manifest.html
[dependencies] [dependencies]
anyhow = "1.0.68" anyhow = "1.0.68"
bytesize = "1.1.0" bytesize = "2.0.1"
chrono = "0.4.23" chrono = "0.4.23"
clap = { version = "4.0.32", features = ["derive"] } clap = { version = "4.0.32", features = ["derive"] }
colored = "2.0.0" dirs = "6.0.0"
dashmap = { version = "5.4.0", features = ["rayon"] } globwalk = "0.9.1"
globwalk = "0.8.1" gxhash = { version = "3.4.1", default-features = false }
gxhash = "3.4.1" indicatif = { version = "0.18.0", features = ["rayon"] }
indicatif = { version = "0.17.2", features = ["rayon"] } memmap2 = "0.9.7"
itertools = "0.10.5" open = "5.4.0"
memmap2 = "0.5.8"
pathdiff = "0.2.1" pathdiff = "0.2.1"
prettytable-rs = "0.10.0" prettytable-rs = "0.10.0"
ratatui = "0.29.0" ratatui = "0.30.2"
rayon = "1.6.1" rayon = "1.6.1"
serde = { version = "1.0.192", features = ["derive"] } unicode-segmentation = "1.12.0"
serde_json = "1.0.108"
tempfile = "3.20.0"
threadpool = "1.8.1"
unicode-segmentation = "1.10.0"
uuid = { version = "1.17.0", features = ["v4"] }
[profile.release] [profile.release]
strip = true strip = true
opt-level = 3
lto = "thin"
debug = false
codegen-units = 1
# generated by 'cargo dist init' # generated by 'cargo dist init'
[profile.dist] [profile.dist]
inherits = "release" inherits = "release"
lto = "thin"
[workspace.metadata.dist] [workspace.metadata.dist]
rust-toolchain-version = "1.78.0" rust-toolchain-version = "1.87.0"
ci = ["github"] ci = ["github"]
targets = ["x86_64-unknown-linux-gnu", "x86_64-apple-darwin", "x86_64-pc-windows-msvc", "aarch64-apple-darwin"] targets = [
"x86_64-unknown-linux-gnu",
"x86_64-apple-darwin",
"x86_64-pc-windows-msvc",
"aarch64-apple-darwin",
]
cargo-dist-version = "0.0.7" cargo-dist-version = "0.0.7"
[dev-dependencies]
tempfile = "3.20.0"
rand = "0.9.1"

228
README.md
View File

@@ -7,21 +7,25 @@
## Usage ## Usage
```bash ```bash
find,filter and delete duplicate files
Usage: deduplicator [OPTIONS] [scan_dir_path] Usage: deduplicator [OPTIONS] [scan_dir_path]
Arguments: Arguments:
[scan_dir_path] Run Deduplicator on dir different from pwd (e.g., ~/Pictures ) [scan_dir_path] Run Deduplicator on dir different from pwd (e.g., ~/Pictures )
Options: Options:
-T, --exclude-types <EXCLUDE_TYPES> Exclude Filetypes [default = none]
-t, --types <TYPES> Filetypes to deduplicate [default = all] -t, --types <TYPES> Filetypes to deduplicate [default = all]
-i, --interactive Delete files interactively -i, --interactive Delete files interactively
-s, --min-size <MIN_SIZE> Minimum filesize of duplicates to scan (e.g., 100B/1K/2M/3G/4T) [default: 1b] -m, --min-size <MIN_SIZE> Minimum filesize of duplicates to scan (e.g., 100B/1K/2M/3G/4T) [default: 1b]
-d, --max-depth <MAX_DEPTH> Max Depth to scan while looking for duplicates -D, --max-depth <MAX_DEPTH> Max Depth to scan while looking for duplicates
--min-depth <MIN_DEPTH> Min Depth to scan while looking for duplicates -d, --min-depth <MIN_DEPTH> Min Depth to scan while looking for duplicates
-f, --follow-links Follow links while scanning directories -f, --follow-links Follow links while scanning directories
-h, --help Print help information -s, --strict Guarantees that two files are duplicate (performs a full hash)
-V, --version Print version information -p, --progress Show Progress spinners & metrics
--json -h, --help Print help
-V, --version Print version
``` ```
### Examples ### Examples
@@ -29,6 +33,9 @@ Options:
# Scan for duplicates recursively from the current dir, only look for png, jpg & pdf file types & interactively delete files # Scan for duplicates recursively from the current dir, only look for png, jpg & pdf file types & interactively delete files
deduplicator -t pdf,jpg,png -i deduplicator -t pdf,jpg,png -i
# Scan for duplicates recursively from current dir, exclude png and jpg file types
deduplicator -T jpg,png
# Scan for duplicates recursively from the ~/Pictures dir, only look for png, jpeg, jpg & pdf file types & interactively delete files # Scan for duplicates recursively from the ~/Pictures dir, only look for png, jpeg, jpg & pdf file types & interactively delete files
deduplicator ~/Pictures/ -t png,jpeg,jpg,pdf -i deduplicator ~/Pictures/ -t png,jpeg,jpg,pdf -i
@@ -42,94 +49,165 @@ deduplicator ~/.config --follow-links
deduplicator ~/Media --min-size 100mb deduplicator ~/Media --min-size 100mb
``` ```
## Demo
![demo](https://github.com/user-attachments/assets/bdb95831-542d-4902-a458-4e0f5d171a33)
## Installation ## Installation
Currently, you can only install deduplicator using cargo package manager.
### Cargo Install ### Cargo
> GxHash relies on aes hardware acceleration, so please set `RUSTFLAGS` to `"-C target-feature=+aes"` or `"-C target-cpu=native"` before
> installing.
#### Stable #### install from crates.io (stable)
> [!WARNING] Note from GxHash: GxHash relies on aes hardware acceleration, you must make sure the aes feature is enabled when building (otherwise it won't build). This can be done by setting the RUSTFLAGS environment variable to -C target-feature=+aes or -C target-cpu=native (the latter should work if your CPU is properly recognized by rustc, which is the case most of the time).
> please install version `0.2.1` if you are unable to install `0.2.2`
```bash ```bash
$ RUSTFLAGS="-C target-cpu=native" cargo install deduplicator $ RUSTFLAGS="-C target-cpu=native" cargo install deduplicator
# or
$ RUSTFLAGS="-C target-feature=+aes,+sse2" cargo install deduplicator
``` ```
> [!] #### install from git (nightly)
#### Nightly
if you'd like to install with nightly features, you can use
```bash ```bash
$ cargo install --git https://github.com/sreedevk/deduplicator $ RUSTFLAGS="-C target-cpu=native" cargo install deduplicator --git https://github.com/sreedevk/deduplicator
```
Please note that if you use a version manager to install rust (like asdf), you need to reshim (`asdf reshim rust`).
### Linux (Pre-built Binary) # or
you can download the pre-built binary from the [Releases](https://github.com/sreedevk/deduplicator/releases) page. $ RUSTFLAGS="-C target-feature=+aes,+sse2" cargo install --git https://github.com/sreedevk/deduplicator
download the `deduplicator-x86_64-unknown-linux-gnu.tar.gz` for linux. Once you have the tarball file with the executable,
you can follow these steps to install:
```bash
$ tar -zxvf deduplicator-x86_64-unknown-linux-gnu.tar.gz
$ sudo mv deduplicator /usr/bin/
``` ```
### Mac OS (Pre-built Binary) ### Manual Installation
- Download the right pre-compiled binary archive for your platform from [github release page](https://github.com/sreedevk/deduplicator/releases/tag/latest).
you can download the pre-build binary from the [Releases](https://github.com/sreedevk/deduplicator/releases) page. - Decompress it using `tar -zxvf <archive>.tar.gz`
download the `deduplicator-x86_64-apple-darwin.tar.gz` tarball for mac os. Once you have the tarball file with the executable, you can follow these steps to install: - Move it to a directory included in `$PATH`.
- ideally `/usr/local/bin/`.
```bash
$ tar -zxvf deduplicator-x86_64-unknown-linux-gnu.tar.gz
$ sudo mv deduplicator /usr/bin/
```
### Windows (Pre-built Binary)
you can download the pre-build binary from the [Releases](https://github.com/sreedevk/deduplicator/releases) page.
download the `deduplicator-x86_64-pc-windows-msvc.zip` zip file for windows. unzip the `zip` file & move the `deduplicator.exe` to a location in the PATH system environment variable.
Note: If you Run into an msvc error, please install MSCV from [here](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist?view=msvc-170)
## Performance ## Performance
Deduplicator uses size comparison and [GxHash](https://docs.rs/gxhash/latest/gxhash/) to quickly check a large number of files to find duplicates. its also heavily parallelized. The default behavior of deduplicator is to only hash the first page (4K) of the file. This is to ensure that performance is the default priority. You can modify this behavior by using the `--strict` flag which will hash the whole file and ensure that 2 files are indeed duplicates. I'll add benchmarks in future versions.
Deduplicator uses size comparison and fxhash (a non non-cryptographic hashing algo) to quickly scan through large number of files to find duplicates. its also highly parallel (uses rayon and dashmap). I was able to scan through 120GB of files (Videos, PDFs, Images) in ~300ms. checkout the benchmarks ### Benchmarks
I've used hyperfine to run deduplicator on files generated by the rake file at `rakelib/benchmark.rake`. The Benchmarking accuracy can further be improved by isolating runs inside restricted docker containers. I'll include that in the future. For now, here's the hyperfine output on my i7-12800H laptop with 32G of RAM.
## benchmarks
| Command | Dirsize | Filecount | Mean [ms] | Min [ms] | Max [ms] | Relative |
|:---|:---|---:|---:|---:|---:|---:|
| `deduplicator ~/Data/tmp` | (~120G) | 721 files | 33.5 ± 28.6 | 25.3 | 151.5 | 1.87 ± 1.60 |
| `deduplicator ~/Data/books` | (~8.6G) | 1419 files | 24.5 ± 1.0 | 22.9 | 28.1 | 1.37 ± 0.08 |
| `deduplicator ~/Data/books --min-size 10M` | (~8.6G) | 1419 files | 17.9 ± 0.7 | 16.8 | 20.0 | 1.00 |
| `deduplicator ~/Data/ --types pdf,jpg,png,jpeg` | (~290G) | 104222 files | 1207.2 ± 37.0 | 1172.2 | 1287.7 | 67.27 ± 3.33 |
* The last entry is lower because of the number of files deduplicator had to go through (~660895 Files). The average size of the files rarely affect the performance of deduplicator.
These benchmarks were run using [hyperfine](https://github.com/sharkdp/hyperfine). Here are the specs of the machine used to benchmark deduplicator:
#### Fewer Large Files
``` ```
OS: Arch Linux x86_64 # hyperfine -N --warmup 80 './target/release/deduplicator bench_artifacts'
Host: Precision 5540 Benchmark 1: ./target/release/deduplicator bench_artifacts
Kernel: 5.15.89-1-lts Time (mean ± σ): 2.2 ms ± 0.4 ms [User: 2.2 ms, System: 4.4 ms]
Uptime: 4 hours, 44 mins Range (min … max): 1.3 ms … 7.1 ms 1522 runs
Shell: zsh 5.9
Terminal: kitty dust 'bench_artifacts'
CPU: Intel i9-9880H (16) @ 4.800GHz
GPU: NVIDIA Quadro T2000 Mobile / Max-Q 54M ┌── file_0_fwds.bin │████ │ 2%
GPU: Intel CoffeeLake-H GT2 [UHD Graphics 630] 122M ├── file_1_fwds.bin │████████ │ 5%
Memory: 31731MiB (~32GiB) 390M ├── file_0_fwdcbss.bin│██████████████████████████ │ 15%
390M ├── file_0_fwscas.bin │██████████████████████████ │ 15%
390M ├── file_0_fwss.bin │██████████████████████████ │ 15%
390M ├── file_1_fwdcbss.bin│██████████████████████████ │ 15%
390M ├── file_1_fwscas.bin │██████████████████████████ │ 15%
390M ├── file_1_fwss.bin │██████████████████████████ │ 15%
2.5G ┌─┴ bench_artifacts │██████████████████████████████████████████████████████████████████ │ 100%
``` ```
## Screenshots #### Many Small Files
```
# hyperfine --warmup 20 './target/release/deduplicator bench_artifacts'
Benchmark 1: ./target/release/deduplicator bench_artifacts
Time (mean ± σ): 40.1 ms ± 2.3 ms [User: 251.0 ms, System: 277.3 ms]
Range (min … max): 35.0 ms … 45.9 ms 72 runs
![](https://user-images.githubusercontent.com/36154121/213618143-e5182e39-731e-4817-87dd-1a6a0f38a449.gif) dust 'bench_artifacts'
3.9M ┌── file_992_fwscas.bin │█ │ 0%
3.9M ├── file_992_fwss.bin │█ │ 0%
3.9M ├── file_993_fwdcbss.bin│█ │ 0%
3.9M ├── file_993_fwscas.bin │█ │ 0%
3.9M ├── file_993_fwss.bin │█ │ 0%
3.9M ├── file_994_fwdcbss.bin│█ │ 0%
3.9M ├── file_994_fwscas.bin │█ │ 0%
3.9M ├── file_994_fwss.bin │█ │ 0%
3.9M ├── file_995_fwdcbss.bin│█ │ 0%
3.9M ├── file_995_fwscas.bin │█ │ 0%
3.9M ├── file_995_fwss.bin │█ │ 0%
3.9M ├── file_996_fwdcbss.bin│█ │ 0%
3.9M ├── file_996_fwscas.bin │█ │ 0%
3.9M ├── file_996_fwss.bin │█ │ 0%
3.9M ├── file_997_fwdcbss.bin│█ │ 0%
3.9M ├── file_997_fwscas.bin │█ │ 0%
3.9M ├── file_997_fwss.bin │█ │ 0%
3.9M ├── file_998_fwdcbss.bin│█ │ 0%
3.9M ├── file_998_fwscas.bin │█ │ 0%
3.9M ├── file_998_fwss.bin │█ │ 0%
3.9M ├── file_999_fwdcbss.bin│█ │ 0%
3.9M ├── file_999_fwscas.bin │█ │ 0%
3.9M ├── file_999_fwss.bin │█ │ 0%
3.9M ├── file_99_fwdcbss.bin │█ │ 0%
3.9M ├── file_99_fwscas.bin │█ │ 0%
3.9M ├── file_99_fwss.bin │█ │ 0%
3.9M ├── file_9_fwdcbss.bin │█ │ 0%
3.9M ├── file_9_fwscas.bin │█ │ 0%
3.9M ├── file_9_fwss.bin │█ │ 0%
11G ┌─┴ bench_artifacts │████████████████████████████████████████████████████████████████ │ 100%
```
## Roadmap ## proposed
- Tree format output for duplicate file listing - [ ] parallelization
- GUI - [ ] scanning + processing sw + processing hw + formatting + printing
- Packages for different operating system repositories (currently only installable via cargo) - [ ] user supplied cache file path for faster re-runs
- TUI: Improve Key Handling - [ ] hardlinks / symlinks support
- [ ] max file path size should use the last set of duplicates
- [ ] add more unit tests
- [ ] test against different filesystems
- [ ] test against different file name encodings
- [ ] restore json output (was removed in 0.3 due to quality issues)
- [ ] fix memory leak on very large filesystems
- [ ] maybe use a bloom filter
- [ ] reduce FileInfo size
- [ ] tui
- [ ] change the default hashing method to include the first & last page of a file (8K)
- [ ] provide option to localize duplicate detection to arbitrary levels relative to current directory
- [ ] localize file meta store locks to sub path levels to avoid global lock contention from multiple threads.
- [ ] bulk operations
- [ ] --keep-latest
- [ ] --keep-oldest
- [ ] --keep-last-modified
- [ ] --keep-first-modified
- [ ] fix: partial hash collision - a file full of null bytes ("\0") and an empty file. This is a known trade off in gxhash.
- [ ] include initial pages and final pages of the file
- [ ] append the offset between the last initial page hashed and the first final page hashed in the content passed to the hasher.
- [ ] potential optimizations
- [ ] lookup the memory efficiency gains if instead of directly inserting into a hashmap, deduplicator looksup the file in a bloom filter. this way, the duplicate store hashmap does
not require to be locked for every single file. The bloom filter can be stored in an atomically updateable type to improve performance as well.
## v0.3.2
- [x] fix: single file groups are printed to screen
- [x] fix: --exclude-types and --types flag behave identically.
## v0.3.1
- [x] parallelization
- [x] (scanning + processing sw + processing hw) & formatting & printing
- [x] remove formatting step and write directly to stdout
- [x] simplify output to improve performance
- [x] increase the number of pages hashed in partial hashing
- [x] updated dependencies
- [x] fix: full file hash collision between a file full of null bytes ("\0") and an empty file. This is a known trade off in gxhash.
- [x] appending the file size at the end of content before hashing.
- [x] created an automated release pipeline
## v0.3.0
- [x] parallelization
- [x] (scanning) + (processing sw & processing hw & formatting & printing)
- [x] reduce cloning values on the heap
- [x] add a partial hashing mode (--strict)
- [x] add unit tests
- [x] add silent mode
- [x] update documentation
- [x] remove color output
- [x] progress bar improvements
- [x] use progress bar groups
- [x] remove broken json rendering
- [x] add benchmarks

1
Rakefile Normal file
View File

@@ -0,0 +1 @@
# tasks are contained in the `rakelib` directory

102
rakelib/benchmark.rake Normal file
View File

@@ -0,0 +1,102 @@
require 'tempfile'
require 'fileutils'
require 'securerandom'
namespace :benchmark do
file 'target/release/deduplicator' do
sh "cargo build --release"
end
task :few_large_files => 'target/release/deduplicator' do
# Benchmark 1: ./target/release/deduplicator bench_artifacts
# Time (mean ± σ): 2.5 ms ± 0.6 ms [User: 2.4 ms, System: 4.8 ms]
# Range (min … max): 1.8 ms … 9.7 ms 1474 runs
root = "bench_artifacts"
FileUtils.rm_rf(root)
Dir.mkdir(root)
# files with same size
puts "generating files of same size ..."
2.times.map do |i|
File.open(File.join(root, "file_#{i}_fwss.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * 100_000))
end
end
# files with different sizes
puts "generating files of different sizes ..."
2.times.map do |i|
File.open(File.join(root, "file_#{i}_fwds.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * (rand * 100_000).ceil))
end
end
# files with same content & size
puts "generating files of same content and sizes ..."
2.times.each do |i|
File.open(File.join(root, "file_#{i}_fwscas.bin"), 'wb') do |f|
f.write("\0" * (4096 * 100_000))
end
end
# files with different content but same size
puts "generating files of different content but same sizes ..."
2.times.each do |i|
File.open(File.join(root, "file_#{i}_fwdcbss.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * 100_000))
end
end
sh("hyperfine -N --warmup 80 './target/release/deduplicator #{root}'")
sh("dust '#{root}'")
FileUtils.rm_rf(root)
end
task :many_small_files => 'target/release/deduplicator' do
# Benchmark 1: ./target/release/deduplicator bench_artifacts
# Time (mean ± σ): 10.6 ms ± 1.0 ms [User: 20.0 ms, System: 22.5 ms]
# Range (min … max): 8.4 ms … 14.2 ms 235 runs
root = "bench_artifacts"
Dir.mkdir(root)
# files with same size
puts "generating 1000 files of the same size ... "
1000.times.each do |i|
File.open(File.join(root, "file_#{i}_fwss.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * 1000))
end
end
# files with different sizes
puts "generating 1000 files of different sizes ... "
1000.times.each do |i|
File.open(File.join(root, "file_#{i}_fwds.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * (rand * 100).ceil))
end
end
# files with same content & size
puts "generating files of same content and sizes ..."
1000.times.each do |i|
File.open(File.join(root, "file_#{i}_fwscas.bin"), 'wb') do |f|
f.write("\0" * (4096 * 1000))
end
end
# files with different content but same size
puts "generating files of different content but same sizes ..."
1000.times.each do |i|
File.open(File.join(root, "file_#{i}_fwdcbss.bin"), 'wb') do |f|
f.write(SecureRandom.bytes(4096 * 1000))
end
end
sh("hyperfine --warmup 20 './target/release/deduplicator #{root}'")
sh("dust '#{root}'")
FileUtils.rm_rf(root)
end
end

View File

@@ -1,72 +0,0 @@
use std::sync::mpsc::channel;
use std::sync::Arc;
use anyhow::{anyhow, Result};
use threadpool::ThreadPool;
use crate::params::Params;
use crate::server::{Message, Server};
use crate::tui::Tui;
pub struct App {
tpool: ThreadPool,
server: Arc<Server>,
app_opts: Arc<Params>,
}
impl App {
pub fn new(app_opts: Arc<Params>) -> Self {
Self {
tpool: ThreadPool::new(8),
server: Arc::new(Server::new().expect("server init failed")),
app_opts,
}
}
pub fn start(&self) -> Result<()> {
let (server_tx, server_rx) = channel::<Message>();
let (app_tx, app_rx) = channel::<Message>();
let server_ptr = self.server.clone();
let tui_server_ptr = self.server.clone();
let mut ui = Tui::new(app_tx.clone(), tui_server_ptr);
let root_dir = self
.app_opts
.get_directory()?
.into_os_string()
.into_string()
.map_err(|_| anyhow!("path to str conv failed"))?
.into_boxed_str();
self.tpool.execute(move || {
server_ptr.start(server_rx).expect("server init failed");
});
self.tpool.execute(move || {
ui.start().expect("ui init failed");
});
self.tpool.execute(move || loop {
match app_rx.try_recv() {
Ok(Message::Exit) => {
server_tx
.send(Message::Exit)
.expect("message passing to app from ui failed.");
break;
}
Ok(Message::AddScanDirectory(dir)) => {
server_tx
.send(Message::AddScanDirectory(dir))
.expect("message passing to server failed.");
}
_ => continue,
}
});
app_tx.send(Message::AddScanDirectory(root_dir))?;
self.tpool.join();
Ok(())
}
}

249
src/cache.rs Normal file
View File

@@ -0,0 +1,249 @@
use std::collections::HashMap;
use std::fs;
use std::path::{Path, PathBuf};
use std::time::{SystemTime, UNIX_EPOCH};
const MAGIC: &[u8; 4] = b"DDUP";
const VERSION: u8 = 1;
const RECORD_FIXED_LEN: usize = 45;
pub const CACHE_TTL_DAYS: u64 = 30;
pub struct CacheEntry {
pub size: u64,
pub mtime: u64,
pub strict: bool,
pub hash: u128,
pub last_seen: u64,
}
pub struct Cache {
entries: HashMap<PathBuf, CacheEntry>,
enabled: bool,
}
pub fn mtime_nanos(modified: SystemTime) -> u64 {
modified
.duration_since(UNIX_EPOCH)
.map(|d| d.as_nanos() as u64)
.unwrap_or(0)
}
impl Cache {
pub fn disabled() -> Self {
Self {
entries: HashMap::new(),
enabled: false,
}
}
fn empty_enabled() -> Self {
Self {
entries: HashMap::new(),
enabled: true,
}
}
pub fn load(path: &Path) -> Self {
match fs::read(path) {
Ok(bytes) => Self::parse(&bytes).unwrap_or_else(Self::empty_enabled),
Err(_) => Self::empty_enabled(),
}
}
fn parse(bytes: &[u8]) -> Option<Self> {
if bytes.len() < 5 || &bytes[0..4] != MAGIC || bytes[4] != VERSION {
return None;
}
let mut entries = HashMap::new();
let mut cur = &bytes[5..];
while cur.len() >= RECORD_FIXED_LEN {
let hash = u128::from_le_bytes(cur[0..16].try_into().ok()?);
let size = u64::from_le_bytes(cur[16..24].try_into().ok()?);
let mtime = u64::from_le_bytes(cur[24..32].try_into().ok()?);
let last_seen = u64::from_le_bytes(cur[32..40].try_into().ok()?);
let strict = cur[40] != 0;
let path_len = u32::from_le_bytes(cur[41..45].try_into().ok()?) as usize;
let rest = &cur[RECORD_FIXED_LEN..];
if rest.len() < path_len {
break;
}
match std::str::from_utf8(&rest[..path_len]) {
Ok(s) => {
entries.insert(
PathBuf::from(s),
CacheEntry {
size,
mtime,
strict,
hash,
last_seen,
},
);
}
Err(_) => {}
}
cur = &rest[path_len..];
}
Some(Self {
entries,
enabled: true,
})
}
pub fn lookup(&self, path: &Path, size: u64, mtime: u64, strict: bool) -> Option<u128> {
let entry = self.entries.get(path)?;
match entry.size == size && entry.mtime == mtime && entry.strict == strict {
true => Some(entry.hash),
false => None,
}
}
pub fn record(
&mut self,
path: &Path,
size: u64,
mtime: u64,
strict: bool,
hash: u128,
now_secs: u64,
) {
if path.to_str().is_none() {
return;
}
self.entries.insert(
path.to_path_buf(),
CacheEntry {
size,
mtime,
strict,
hash,
last_seen: now_secs,
},
);
}
pub fn save(&self, path: &Path, ttl_secs: u64, now_secs: u64) {
if !self.enabled {
return;
}
let mut buf: Vec<u8> = Vec::new();
buf.extend_from_slice(MAGIC);
buf.push(VERSION);
for (entry_path, entry) in &self.entries {
if now_secs.saturating_sub(entry.last_seen) > ttl_secs {
continue;
}
let path_str = match entry_path.to_str() {
Some(s) => s,
None => continue,
};
buf.extend_from_slice(&entry.hash.to_le_bytes());
buf.extend_from_slice(&entry.size.to_le_bytes());
buf.extend_from_slice(&entry.mtime.to_le_bytes());
buf.extend_from_slice(&entry.last_seen.to_le_bytes());
buf.push(entry.strict as u8);
buf.extend_from_slice(&(path_str.len() as u32).to_le_bytes());
buf.extend_from_slice(path_str.as_bytes());
}
if let Some(parent) = path.parent() {
let _ = fs::create_dir_all(parent);
}
if let Err(err) = fs::write(path, &buf) {
eprintln!("warning: failed to write cache to {}: {err}", path.display());
}
}
pub fn default_path() -> Option<PathBuf> {
dirs::cache_dir().map(|dir| dir.join("deduplicator").join("cache.bin"))
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::path::Path;
use std::time::{Duration, UNIX_EPOCH};
use tempfile::TempDir;
#[test]
fn lookup_hits_on_exact_match() {
let mut c = Cache::disabled();
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), Some(42));
}
#[test]
fn lookup_misses_on_any_field_change_or_absent() {
let mut c = Cache::disabled();
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
assert_eq!(c.lookup(Path::new("/a/b"), 101, 200, false), None);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 201, false), None);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, true), None);
assert_eq!(c.lookup(Path::new("/a/x"), 100, 200, false), None);
}
#[test]
fn save_then_load_roundtrips_entries() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
let mut c = Cache::load(&file);
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
c.record(Path::new("/c/d"), 5, 6, true, 99, 1000);
c.save(&file, 86_400, 1000);
let loaded = Cache::load(&file);
assert_eq!(loaded.lookup(Path::new("/a/b"), 100, 200, false), Some(42));
assert_eq!(loaded.lookup(Path::new("/c/d"), 5, 6, true), Some(99));
}
#[test]
fn load_returns_empty_on_corrupt_file() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
std::fs::write(&file, b"not a valid cache file at all").unwrap();
let c = Cache::load(&file);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), None);
}
#[test]
fn load_returns_empty_on_version_mismatch() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
std::fs::write(&file, b"DDUP\x02").unwrap();
let c = Cache::load(&file);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), None);
}
#[test]
fn save_prunes_entries_older_than_ttl() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
let mut c = Cache::load(&file);
c.record(Path::new("/old"), 1, 2, false, 10, 1000);
c.record(Path::new("/new"), 3, 4, false, 20, 5000);
c.save(&file, 1000, 5000);
let loaded = Cache::load(&file);
assert_eq!(loaded.lookup(Path::new("/old"), 1, 2, false), None);
assert_eq!(loaded.lookup(Path::new("/new"), 3, 4, false), Some(20));
}
#[test]
fn mtime_nanos_converts_systemtime() {
let t = UNIX_EPOCH + Duration::from_nanos(1234);
assert_eq!(mtime_nanos(t), 1234);
}
}

View File

View File

@@ -1,43 +1,85 @@
use anyhow::Result; use anyhow::Result;
use gxhash::GxHasher; use gxhash::gxhash128;
use memmap2::Mmap; use memmap2::Mmap;
use serde::Serialize; use std::{
use std::fs; fs,
use std::hash::Hasher; io::Read,
use std::{fs::Metadata, path::PathBuf}; path::{Path, PathBuf},
time::SystemTime,
};
#[derive(Debug, Clone, Serialize)] #[derive(Debug, Clone)]
pub struct FileInfo { pub struct FileInfo {
pub path: PathBuf, pub path: Box<Path>,
pub hash: Option<String>,
pub size: u64, pub size: u64,
#[serde(skip)] pub modified: SystemTime,
pub filemeta: Metadata,
} }
impl FileInfo { impl FileInfo {
pub fn hash(&self) -> Result<Self> { pub fn hash(&self, seed: i64) -> Result<u128> {
let file = fs::File::open(self.path.clone())?; if self.size == 0 {
return Ok(0u128);
};
let file = fs::File::open(&self.path)?;
let mapper = unsafe { Mmap::map(&file)? }; let mapper = unsafe { Mmap::map(&file)? };
let mut primhasher = GxHasher::default(); let content_hash = mapper
.chunks(4096)
.fold(0u128, |acc, chunk: &[u8]| acc ^ gxhash128(chunk, seed));
mapper // NOTE: avoids collision bw an empty file & a file full of null bytes.
.chunks(1_000_000) Ok(content_hash ^ gxhash128(&self.size.to_ne_bytes(), seed))
.for_each(|chunk| primhasher.write(chunk)); }
Ok(Self { pub fn initpages_hash(&self, seed: i64) -> Result<u128> {
hash: Some(primhasher.finish().to_string()), let mut file = fs::File::open(&self.path)?;
..self.clone() let mut buffer = [0; 16384];
}) let bytes_read = file.read(&mut buffer)?;
Ok(gxhash128(&buffer[..bytes_read], seed))
} }
pub fn new(path: PathBuf) -> Result<Self> { pub fn new(path: PathBuf) -> Result<Self> {
let filemeta = std::fs::metadata(path.clone())?; let filemeta = std::fs::metadata(&path)?;
Ok(Self { Ok(Self {
path, path: path.into_boxed_path(),
filemeta: filemeta.clone(),
hash: None,
size: filemeta.len(), size: filemeta.len(),
modified: filemeta.modified()?,
}) })
} }
} }
#[cfg(test)]
mod test {
use super::*;
use tempfile::TempDir;
use std::fs::File;
use std::io::Write;
use anyhow::Result;
fn generate_null_bytes(size: usize) -> Vec<u8> {
(0..size).map(|_| 0).collect::<Vec<u8>>()
}
#[test]
fn hash_differentiates_between_a_file_of_null_bytes_vs_an_empty_file() -> Result<()> {
let root = TempDir::new()?;
let empty_file_name = root.path().join("empty_file.bin");
File::create_new(&empty_file_name)?;
let file_with_null_bytes_name = root.path().join("file_with_null_bytes.bin");
let mut file_with_null_bytes = File::create_new(&file_with_null_bytes_name)?;
file_with_null_bytes.write_all(&generate_null_bytes(1000 * 4096))?;
let empty_file_info = FileInfo::new(empty_file_name)?;
let file_with_empty_bytes_info = FileInfo::new(file_with_null_bytes_name)?;
let seed: i64 = 246910456374;
assert_ne!(empty_file_info.hash(seed)?, file_with_empty_bytes_info.hash(seed)?);
Ok(())
}
}

View File

@@ -1,33 +1,23 @@
pub struct Formatter; use crate::{fileinfo::FileInfo, params::Params, pipeline::DedupReport};
use crate::fileinfo::FileInfo;
use crate::params::Params;
use anyhow::Result; use anyhow::Result;
use chrono::{DateTime, Utc}; use chrono::{DateTime, Utc};
use colored::Colorize;
use dashmap::DashMap;
use indicatif::{
ParallelProgressIterator, ProgressBar, ProgressFinish, ProgressIterator, ProgressStyle,
};
use pathdiff::diff_paths; use pathdiff::diff_paths;
use prettytable::{format, row, Table};
use rayon::prelude::*; use rayon::prelude::*;
use std::borrow::Cow;
use std::path::PathBuf; use std::path::PathBuf;
use std::time::Duration;
const YELLOW: &str = "\x1b[33m";
const RESET: &str = "\x1b[0m";
pub struct Formatter;
impl Formatter { impl Formatter {
pub fn human_path( pub fn human_path(file: &FileInfo, aargs: &Params, max_path_length: usize) -> Result<String> {
file: &FileInfo, let base_directory: PathBuf = aargs.get_directory()?;
app_args: &Params, let relative_path = diff_paths(&file.path, base_directory).unwrap_or_default();
min_path_length: usize,
) -> Result<String> {
let base_directory: PathBuf = app_args.get_directory()?;
let relative_path = diff_paths(file.path.clone(), base_directory).unwrap_or_default();
let formatted_path = format!( let formatted_path = format!(
"{:<0width$}", "{:<0width$}",
relative_path.to_str().unwrap_or_default().to_string(), relative_path.to_str().unwrap_or_default().to_string(),
width = min_path_length width = max_path_length
); );
Ok(formatted_path) Ok(formatted_path)
@@ -38,97 +28,43 @@ impl Formatter {
} }
pub fn human_mtime(file: &FileInfo) -> Result<String> { pub fn human_mtime(file: &FileInfo) -> Result<String> {
let modified_time: DateTime<Utc> = file.filemeta.modified()?.into(); let modified_time: DateTime<Utc> = file.modified.into();
Ok(modified_time.format("%Y-%m-%d %H:%M:%S").to_string()) Ok(modified_time.format("%Y-%m-%d %H:%M:%S").to_string())
} }
pub fn generate_table(raw: Vec<FileInfo>, app_args: &Params) -> Result<Table> { pub fn print(report: &DedupReport, aargs: &Params) {
let basepath_length = app_args.get_directory()?.to_str().unwrap_or_default().len(); print!("{}", "\n".repeat(if aargs.progress { 2 } else { 1 }));
let max_filepath_length = raw
.iter()
.map(|file| file.path.to_str().unwrap_or_default().len())
.max()
.unwrap_or_default();
let min_path_length = if max_filepath_length > basepath_length { if report.groups.is_empty() {
max_filepath_length - basepath_length println!("No duplicates found matching your search criteria.");
return;
}
report.groups.par_iter().for_each(|group| {
let mut ostring = format!("{}{:32x}{}\n", YELLOW, group.hash, RESET);
let subfields = group
.files
.par_iter()
.enumerate()
.map(|(i, finfo)| {
let nodechar = if i == group.files.len() - 1 {
"└─"
} else { } else {
0 "├─"
}; };
format!(
"{}\t{}\t{}\t{}\n",
nodechar,
Self::human_path(finfo, aargs, report.max_path_len)
.expect("path formatting failed."),
Self::human_filesize(finfo).expect("filesize formatting failed."),
Self::human_mtime(finfo).expect("modified time formatting failed.")
)
})
.collect::<String>();
let progress_style = ProgressStyle::with_template( ostring.push_str(&subfields);
"[{elapsed_precise}] {bar:40.cyan/blue} {pos:>7}/{len:7} {msg}", println!("{ostring}");
)?;
let progress_bar = ProgressBar::new(raw.len() as u64);
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("reconciling data");
let duplicates_table: DashMap<String, Vec<FileInfo>> = DashMap::new();
raw.into_par_iter()
.progress_with(progress_bar)
.with_finish(ProgressFinish::WithMessage(Cow::from("data reconciled")))
.map(|file| file.hash())
.filter_map(Result::ok)
.for_each(|file| {
duplicates_table
.entry(file.hash.clone().unwrap_or_default())
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file]);
}); });
let mut output_table = Table::new();
output_table.set_titles(row!["hash", "duplicates"]);
let progress_style = ProgressStyle::with_template(
"[{elapsed_precise}] {bar:40.cyan/blue} {pos:>7}/{len:7} {msg}",
)?;
let progress_bar = ProgressBar::new(duplicates_table.len() as u64);
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("generating output");
duplicates_table
.into_iter()
.progress_with(progress_bar)
.with_finish(ProgressFinish::WithMessage(Cow::from("output generated")))
.for_each(|(hash, group)| {
let mut inner_table = Table::new();
inner_table.set_format(*format::consts::FORMAT_NO_BORDER_LINE_SEPARATOR);
group.iter().for_each(|file| {
inner_table.add_row(row![
Self::human_path(file, app_args, min_path_length)
.unwrap_or_default()
.blue(),
Self::human_filesize(file).unwrap_or_default().red(),
Self::human_mtime(file).unwrap_or_default().yellow()
]);
});
output_table.add_row(row![hash.green(), inner_table]);
});
Ok(output_table)
}
pub fn print(raw: Vec<FileInfo>, app_args: &Params) -> Result<()> {
if raw.is_empty() {
println!(
"\n\n{}\n",
"No duplicates found matching your search criteria.".green()
);
return Ok(());
}
if app_args.json {
let output_json = serde_json::to_string_pretty(&raw)?;
println!("{}", output_json);
} else {
let output_table = Self::generate_table(raw, app_args)?;
output_table.printstd();
}
Ok(())
} }
} }

View File

@@ -1,16 +1,44 @@
use crate::formatter::Formatter; use crate::{fileinfo::FileInfo, formatter::Formatter, params::Params, pipeline::DuplicateGroup};
use crate::{fileinfo::FileInfo, params::Params};
use anyhow::Result; use anyhow::Result;
use colored::Colorize;
use dashmap::DashMap;
use indicatif::{ParallelProgressIterator, ProgressBar, ProgressFinish, ProgressStyle};
use prettytable::{format, row, Table}; use prettytable::{format, row, Table};
use rayon::prelude::*; use std::io::{self, Write};
use std::{ use unicode_segmentation::UnicodeSegmentation;
borrow::Cow,
io::{self, Write}, pub struct Interactive;
time::Duration,
}; impl Interactive {
pub fn init(groups: &[DuplicateGroup], app_args: &Params) -> Result<()> {
if groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
}
groups.iter().enumerate().for_each(|(gindex, group)| {
let files = &group.files;
let mut itable = Table::new();
itable.set_format(*format::consts::FORMAT_NO_BORDER_LINE_SEPARATOR);
itable.set_titles(row!["index", "filename", "size", "updated_at"]);
let max_path_size = files
.iter()
.map(|f| f.path.to_string_lossy().graphemes(true).count())
.max()
.unwrap_or_default();
files.iter().enumerate().for_each(|(index, file)| {
itable.add_row(row![
index,
Formatter::human_path(file, app_args, max_path_size).unwrap_or_default(),
Formatter::human_filesize(file).unwrap_or_default(),
Formatter::human_mtime(file).unwrap_or_default()
]);
});
Self::process_group_action(files, gindex, groups.len(), itable);
});
Ok(())
}
pub fn scan_group_confirmation() -> Result<bool> { pub fn scan_group_confirmation() -> Result<bool> {
print!("\nconfirm? [y/N]: "); print!("\nconfirm? [y/N]: ");
@@ -36,67 +64,6 @@ pub fn scan_group_instruction() -> Result<String> {
Ok(user_input) Ok(user_input)
} }
pub fn init(result: Vec<FileInfo>, app_args: &Params) -> Result<()> {
let basepath_length = app_args.get_directory()?.to_str().unwrap_or_default().len();
let max_filepath_length = result
.iter()
.map(|file| file.path.to_str().unwrap_or_default().len())
.max()
.unwrap_or_default();
let min_path_length = if max_filepath_length > basepath_length {
max_filepath_length - basepath_length
} else {
0
};
let progress_style = ProgressStyle::with_template(
"[{elapsed_precise}] {bar:40.cyan/blue} {pos:>7}/{len:7} {msg}",
)?;
let progress_bar = ProgressBar::new(result.len() as u64);
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("reconciling data");
let duplicates: DashMap<String, Vec<FileInfo>> = DashMap::new();
result
.into_par_iter()
.progress_with(progress_bar)
.with_finish(ProgressFinish::WithMessage(Cow::from("data reconciled")))
.map(|file| file.hash())
.filter_map(Result::ok)
.for_each(|file| {
duplicates
.entry(file.hash.clone().unwrap_or_default())
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file]);
});
duplicates
.clone()
.into_iter()
.enumerate()
.for_each(|(gindex, (_, group))| {
let mut itable = Table::new();
itable.set_format(*format::consts::FORMAT_NO_BORDER_LINE_SEPARATOR);
itable.set_titles(row!["index", "filename", "size", "updated_at"]);
group.iter().enumerate().for_each(|(index, file)| {
itable.add_row(row![
index,
Formatter::human_path(file, app_args, min_path_length)
.unwrap_or_default()
.blue(),
Formatter::human_filesize(file).unwrap_or_default().red(),
Formatter::human_mtime(file).unwrap_or_default().yellow()
]);
});
process_group_action(&group, gindex, duplicates.len(), itable);
});
Ok(())
}
pub fn process_group_action( pub fn process_group_action(
duplicates: &Vec<FileInfo>, duplicates: &Vec<FileInfo>,
dup_index: usize, dup_index: usize,
@@ -105,7 +72,7 @@ pub fn process_group_action(
) { ) {
println!("\nDuplicate Set {} of {}\n", dup_index + 1, dup_size); println!("\nDuplicate Set {} of {}\n", dup_index + 1, dup_size);
table.printstd(); table.printstd();
let files_to_delete = scan_group_instruction().unwrap_or_default(); let files_to_delete = Self::scan_group_instruction().unwrap_or_default();
let parsed_file_indices = files_to_delete let parsed_file_indices = files_to_delete
.trim() .trim()
.split(',') .split(',')
@@ -118,8 +85,8 @@ pub fn process_group_action(
.into_iter() .into_iter()
.any(|index| index > (duplicates.len() - 1)) .any(|index| index > (duplicates.len() - 1))
{ {
println!("{}", "Err: File Index Out of Bounds!".red()); println!("Err: File Index Out of Bounds!");
return process_group_action(duplicates, dup_index, dup_size, table); return Self::process_group_action(duplicates, dup_index, dup_size, table);
} }
print!("{esc}[2J{esc}[1;1H", esc = 27 as char); print!("{esc}[2J{esc}[1;1H", esc = 27 as char);
@@ -132,23 +99,24 @@ pub fn process_group_action(
.into_iter() .into_iter()
.map(|index| duplicates[index].clone()); .map(|index| duplicates[index].clone());
println!("\n{}", "The following files will be deleted:".red()); println!("\nThe following files will be deleted:");
files_to_delete files_to_delete
.clone() .clone()
.enumerate() .enumerate()
.for_each(|(index, file)| { .for_each(|(index, file)| {
println!("{}: {}", index.to_string().blue(), file.path.display()); println!("{}: {}", index, file.path.display());
}); });
match scan_group_confirmation().unwrap() { match Self::scan_group_confirmation().unwrap() {
true => { true => {
files_to_delete.into_iter().for_each(|file| { files_to_delete.into_iter().for_each(|file| {
match std::fs::remove_file(file.path.clone()) { match std::fs::remove_file(file.path.clone()) {
Ok(_) => println!("{}: {}", "DELETED".green(), file.path.display()), Ok(_) => println!("DELETED: {}", file.path.display()),
Err(_) => println!("{}: {}", "FAILED".red(), file.path.display()), Err(_) => println!("FAILED: {}", file.path.display()),
} }
}); });
} }
false => println!("{}", "\nCancelled Delete Operation.".red()), false => println!("\nCancelled Delete Operation."),
}
} }
} }

View File

@@ -1,40 +1,29 @@
mod app; mod cache;
mod fileinfo; mod fileinfo;
mod formatter; mod formatter;
mod interactive; mod interactive;
mod params; mod params;
mod pipeline;
mod processor; mod processor;
mod resolver;
mod scanner; mod scanner;
/* version 2.0 modules*/
mod cli;
mod server;
mod tui; mod tui;
use self::{formatter::Formatter, interactive::Interactive};
use anyhow::Result; use anyhow::Result;
use self::app::App;
use clap::Parser; use clap::Parser;
use params::Params; use params::Params;
use std::sync::Arc;
// use formatter::Formatter;
// use processor::Processor;
// use scanner::Scanner;
fn main() -> Result<()> { fn main() -> Result<()> {
let app_args = Params::parse(); let params = Params::parse();
// let scan_results = Scanner::build(&app_args)?.scan()?; let report = pipeline::run(&params)?;
// let processor = Processor::new(scan_results);
// let results = processor.sizewise()?.hashwise()?;
// match app_args.interactive { match (params.tui, params.keep, params.interactive) {
// false => Formatter::print(results.files, &app_args)?, (true, _, _) => tui::run(report, &params)?,
// true => interactive::init(results.files, &app_args)?, (false, Some(strategy), _) => resolver::run(&report, strategy, params.force, &params)?,
// } (false, None, true) => Interactive::init(&report.groups, &params)?,
(false, None, false) => Formatter::print(&report, &params),
App::new(Arc::new(app_args)) }
.start()
.expect("app init failed.");
Ok(()) Ok(())
} }

View File

@@ -3,9 +3,12 @@ use std::{fs, path::PathBuf};
use anyhow::Result; use anyhow::Result;
use clap::{Parser, ValueHint}; use clap::{Parser, ValueHint};
#[derive(Parser, Debug, Clone)] #[derive(Parser, Debug, Default, Clone)]
#[command(author, version, about, long_about = None)] #[command(author, version, about, long_about = None)]
pub struct Params { pub struct Params {
/// Exclude Filetypes [default = none]
#[arg(short = 'T', long)]
pub exclude_types: Option<String>,
/// Filetypes to deduplicate [default = all] /// Filetypes to deduplicate [default = all]
#[arg(short, long)] #[arg(short, long)]
pub types: Option<String>, pub types: Option<String>,
@@ -15,21 +18,39 @@ pub struct Params {
/// Delete files interactively /// Delete files interactively
#[arg(long, short)] #[arg(long, short)]
pub interactive: bool, pub interactive: bool,
/// Keep one file per duplicate group by this rule and remove the rest
#[arg(long, conflicts_with = "interactive")]
pub keep: Option<crate::resolver::KeepStrategy>,
/// Actually delete the duplicates (without this, --keep only previews)
#[arg(long, visible_alias = "yes", requires = "keep")]
pub force: bool,
/// Minimum filesize of duplicates to scan (e.g., 100B/1K/2M/3G/4T). /// Minimum filesize of duplicates to scan (e.g., 100B/1K/2M/3G/4T).
#[arg(long, short = 's', default_value = "1b")] #[arg(long, short = 'm', default_value = "1b")]
pub min_size: Option<String>, pub min_size: Option<String>,
/// Max Depth to scan while looking for duplicates /// Max Depth to scan while looking for duplicates
#[arg(long, short = 'd')] #[arg(long, short = 'D')]
pub max_depth: Option<usize>, pub max_depth: Option<usize>,
/// Min Depth to scan while looking for duplicates /// Min Depth to scan while looking for duplicates
#[arg(long)] #[arg(long, short = 'd')]
pub min_depth: Option<usize>, pub min_depth: Option<usize>,
/// Follow links while scanning directories /// Follow links while scanning directories
#[arg(long, short)] #[arg(long, short)]
pub follow_links: bool, pub follow_links: bool,
/// print json output /// Guarantees that two files are duplicate (performs a full hash)
#[arg(long, short = 's', default_value = "false")]
pub strict: bool,
/// Show Progress spinners & metrics
#[arg(long, short = 'p', default_value = "false")]
pub progress: bool,
/// Disable the on-disk hash cache
#[arg(long)] #[arg(long)]
pub json: bool, pub no_cache: bool,
/// Use a specific cache file instead of the default location
#[arg(long, value_name = "PATH")]
pub cache_file: Option<PathBuf>,
/// Browse and resolve duplicates in an interactive terminal UI
#[arg(long, conflicts_with_all = ["keep", "interactive"])]
pub tui: bool,
} }
impl Params { impl Params {
@@ -49,8 +70,4 @@ impl Params {
let dir = fs::canonicalize(dir_path)?; let dir = fs::canonicalize(dir_path)?;
Ok(dir) Ok(dir)
} }
pub fn get_types(&self) -> Option<String> {
self.types.clone()
}
} }

182
src/pipeline.rs Normal file
View File

@@ -0,0 +1,182 @@
use anyhow::Result;
use indicatif::{ProgressBar, ProgressStyle};
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use unicode_segmentation::UnicodeSegmentation;
use crate::cache::{mtime_nanos, Cache, CACHE_TTL_DAYS};
use crate::fileinfo::FileInfo;
use crate::params::Params;
use crate::processor;
use crate::scanner::Scanner;
const HASH_SEED: i64 = 0x00DE_D0CA_C4E5_EED1;
fn now_secs() -> u64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
pub struct DuplicateGroup {
pub hash: u128,
pub files: Vec<FileInfo>,
}
pub struct DedupReport {
pub groups: Vec<DuplicateGroup>,
pub max_path_len: usize,
}
pub(crate) fn spinner(enabled: bool, message: &'static str) -> ProgressBar {
let bar = if enabled {
ProgressBar::new_spinner()
} else {
ProgressBar::hidden()
};
let style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")
.expect("valid progress template");
bar.set_style(style);
bar.enable_steady_tick(Duration::from_millis(50));
bar.set_message(message);
bar
}
pub fn run(params: &Params) -> Result<DedupReport> {
let seed = HASH_SEED;
let cache_path = match params.no_cache {
true => None,
false => params.cache_file.clone().or_else(Cache::default_path),
};
let mut cache = match (params.no_cache, &cache_path) {
(false, Some(path)) => Cache::load(path),
_ => Cache::disabled(),
};
let files = Scanner::new(params)?.scan()?;
let size_groups = processor::group_by_size(files, params.progress);
let candidates: Vec<FileInfo> = size_groups
.into_iter()
.filter(|group| group.len() > 1)
.flatten()
.collect();
let max_path_len = candidates
.iter()
.map(|file| file.path.to_string_lossy().graphemes(true).count())
.max()
.unwrap_or(0);
let hashed = processor::hash_candidates(candidates, params.strict, seed, params.progress, &cache);
let now = now_secs();
for (hash, file) in &hashed {
let mtime = mtime_nanos(file.modified);
cache.record(&file.path, file.size, mtime, params.strict, *hash, now);
}
let groups = processor::group_hashed(hashed);
if let Some(path) = &cache_path {
cache.save(path, CACHE_TTL_DAYS * 86_400, now);
}
Ok(DedupReport {
groups,
max_path_len,
})
}
#[cfg(test)]
mod tests {
use super::run;
use crate::params::Params;
use anyhow::Result;
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
#[test]
fn run_finds_a_single_group_of_identical_files() -> Result<()> {
let root = TempDir::new()?;
let duplicate = vec![7u8; 200_000];
for name in ["a.bin", "b.bin"] {
let mut file = File::create_new(root.path().join(name))?;
file.write_all(&duplicate)?;
}
let mut unique = File::create_new(root.path().join("c.bin"))?;
unique.write_all(&vec![9u8; 100_000])?;
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
let report = run(&params)?;
assert_eq!(report.groups.len(), 1);
assert_eq!(report.groups[0].files.len(), 2);
Ok(())
}
}
#[cfg(test)]
mod cache_tests {
use super::run;
use crate::params::Params;
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
fn make_tree(root: &TempDir) {
for name in ["a.bin", "b.bin"] {
let mut f = File::create_new(root.path().join(name)).unwrap();
f.write_all(&vec![7u8; 200_000]).unwrap();
}
let mut u = File::create_new(root.path().join("c.bin")).unwrap();
u.write_all(&vec![9u8; 100_000]).unwrap();
}
#[test]
fn run_twice_with_cache_is_identical_and_writes_the_file() {
let root = TempDir::new().unwrap();
make_tree(&root);
let cache_file = root.path().join("dd.cache");
let params = Params {
dir: Some(root.path().into()),
cache_file: Some(cache_file.clone()),
..Default::default()
};
let first = run(&params).unwrap();
assert!(cache_file.exists());
let second = run(&params).unwrap();
assert_eq!(first.groups.len(), second.groups.len());
assert_eq!(first.groups.len(), 1);
assert_eq!(second.groups[0].files.len(), 2);
}
#[test]
fn no_cache_does_not_write_a_cache_file() {
let root = TempDir::new().unwrap();
make_tree(&root);
let cache_file = root.path().join("dd.cache");
let params = Params {
dir: Some(root.path().into()),
no_cache: true,
cache_file: Some(cache_file.clone()),
..Default::default()
};
run(&params).unwrap();
assert!(!cache_file.exists());
}
}

View File

@@ -1,107 +1,214 @@
use anyhow::Result; use std::collections::HashMap;
use dashmap::DashMap;
use indicatif::{ParallelProgressIterator, ProgressBar, ProgressStyle, ProgressFinish};
use rayon::prelude::{IntoParallelIterator, ParallelIterator}; use rayon::prelude::{IntoParallelIterator, ParallelIterator};
use std::{time::Duration, borrow::Cow};
use crate::cache::{mtime_nanos, Cache};
use crate::fileinfo::FileInfo;
use crate::pipeline::{spinner, DuplicateGroup};
pub fn group_by_size(files: Vec<FileInfo>, progress: bool) -> Vec<Vec<FileInfo>> {
let bar = spinner(progress, "files grouped by size");
let mut buckets: HashMap<u64, Vec<FileInfo>> = HashMap::new();
for file in files {
bar.inc(1);
buckets.entry(file.size).or_default().push(file);
}
bar.finish_with_message("files grouped by size");
buckets.into_values().collect()
}
pub fn hash_candidates(
candidates: Vec<FileInfo>,
strict: bool,
seed: i64,
progress: bool,
cache: &Cache,
) -> Vec<(u128, FileInfo)> {
let bar = spinner(progress, "files grouped by hash");
let hashed: Vec<(u128, FileInfo)> = candidates
.into_par_iter()
.map(|file| {
bar.inc(1);
let mtime = mtime_nanos(file.modified);
let hash = match cache.lookup(&file.path, file.size, mtime, strict) {
Some(cached) => cached,
None => match strict {
true => file.hash(seed).expect("hashing file failed."),
false => file.initpages_hash(seed).expect("hashing file failed."),
},
};
(hash, file)
})
.collect();
bar.finish_with_message("files grouped by hash.");
hashed
}
pub fn group_hashed(hashed: Vec<(u128, FileInfo)>) -> Vec<DuplicateGroup> {
let mut buckets: HashMap<u128, Vec<FileInfo>> = HashMap::new();
for (hash, file) in hashed {
buckets.entry(hash).or_default().push(file);
}
buckets
.into_iter()
.filter(|(_, files)| files.len() > 1)
.map(|(hash, files)| DuplicateGroup { hash, files })
.collect()
}
#[cfg(test)]
mod staged_tests {
use anyhow::Result;
use rand::Rng;
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
use crate::fileinfo::FileInfo; use crate::fileinfo::FileInfo;
#[derive(Debug, Clone)] fn generate_bytes(size: usize) -> Vec<u8> {
pub enum State { let mut rng = rand::rng();
Initial, (0..size).map(|_| rng.random::<u8>()).collect::<Vec<u8>>()
SizeWise,
HashWise,
} }
#[derive(Debug, Clone)] fn write_files(root: &TempDir, specs: Vec<(&str, Vec<u8>)>) -> Result<Vec<FileInfo>> {
pub struct Processor { specs
pub files: Vec<FileInfo>, .into_iter()
pub state: State, .map(|(name, content)| {
} let path = root.path().join(name);
let mut file = File::create_new(&path)?;
impl Processor { file.write_all(&content)?;
pub fn new(files: Vec<FileInfo>) -> Self { FileInfo::new(path)
Self {
files,
state: State::Initial,
}
}
pub fn hashwise(&self) -> Result<Self> {
if self.files.is_empty() {
return Ok(self.clone());
}
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {bar:40.cyan/blue} {pos:>7}/{len:7} {msg}")?;
let progress_bar = ProgressBar::new(self.files.len() as u64);
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("indexing file hashes");
let duplicates_table: DashMap<String, Vec<FileInfo>> = DashMap::new();
self.files
.clone()
.into_par_iter()
.progress_with(progress_bar)
.with_finish(ProgressFinish::WithMessage(Cow::from("indexed files hashes")))
.map(|file| file.hash())
.filter_map(Result::ok)
.for_each(|file| {
duplicates_table
.entry(file.hash.clone().unwrap_or_default())
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file]);
});
let files = duplicates_table
.into_read_only()
.values()
.cloned()
.filter(|subfiles| subfiles.len() > 1)
.flatten()
.collect::<Vec<FileInfo>>();
Ok(Self {
files,
state: State::HashWise,
}) })
.collect()
} }
pub fn sizewise(&self) -> Result<Self> { #[test]
if self.files.is_empty() { fn group_by_size_separates_files_of_different_sizes() -> Result<()> {
return Ok(self.clone()); let root = TempDir::new()?;
let files = write_files(
&root,
vec![
("fileone.bin", generate_bytes(282624)),
("filetwo.bin", generate_bytes(1720320)),
],
)?;
let groups = super::group_by_size(files, false);
assert_eq!(groups.len(), 2);
Ok(())
} }
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {bar:40.cyan/blue} {pos:>7}/{len:7} {msg}")?; #[test]
let progress_bar = ProgressBar::new(self.files.len() as u64); fn group_by_size_buckets_same_size_files_together() -> Result<()> {
progress_bar.set_style(progress_style); let root = TempDir::new()?;
progress_bar.enable_steady_tick(Duration::from_millis(50)); let files = write_files(
progress_bar.set_message("indexing file sizes"); &root,
vec![
("fileone.bin", generate_bytes(282624)),
("filetwo.bin", generate_bytes(282624)),
],
)?;
let duplicates_table: DashMap<u64, Vec<FileInfo>> = DashMap::new(); let groups = super::group_by_size(files, false);
self.files assert_eq!(groups.len(), 1);
.clone() Ok(())
.into_par_iter() }
.progress_with(progress_bar)
.with_finish(ProgressFinish::WithMessage(Cow::from("indexed files sizes")))
.for_each(|file| {
duplicates_table
.entry(file.size)
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file]);
});
let files = duplicates_table #[test]
.into_read_only() fn group_by_hash_fast_mode_matches_identical_init_pages() -> Result<()> {
.values() let root = TempDir::new()?;
.cloned() let shared = generate_bytes(16384);
.filter(|subfiles| subfiles.len() > 1)
.flatten()
.collect::<Vec<FileInfo>>();
Ok(Self { let mut content_x = shared.clone();
files, let mut content_y = shared.clone();
state: State::SizeWise, content_x.extend(generate_bytes(1720320));
}) content_y.extend(generate_bytes(1720320));
let files = write_files(
&root,
vec![("fileone.bin", content_x), ("filetwo.bin", content_y)],
)?;
let groups = super::group_hashed(super::hash_candidates(files, false, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 1);
Ok(())
}
#[test]
fn group_by_hash_strict_mode_rejects_different_tails() -> Result<()> {
let root = TempDir::new()?;
let shared = generate_bytes(16384);
let mut content_x = shared.clone();
let mut content_y = shared.clone();
content_x.extend(generate_bytes(1720320));
content_y.extend(generate_bytes(1720320));
let files = write_files(
&root,
vec![("fileone.bin", content_x), ("filetwo.bin", content_y)],
)?;
let groups = super::group_hashed(super::hash_candidates(files, true, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 0);
Ok(())
}
#[test]
fn group_by_hash_matches_identical_files() -> Result<()> {
let root = TempDir::new()?;
let content = generate_bytes(282624);
let files = write_files(
&root,
vec![
("fileone.bin", content.clone()),
("filetwo.bin", content.clone()),
],
)?;
let groups = super::group_hashed(super::hash_candidates(files, false, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 1);
Ok(())
}
#[test]
fn hash_candidates_uses_cached_hash_when_valid() {
let root = TempDir::new().unwrap();
let path = root.path().join("f.bin");
let mut f = File::create_new(&path).unwrap();
f.write_all(b"real content for cache hit test").unwrap();
let info = FileInfo::new(path.clone()).unwrap();
let mtime = crate::cache::mtime_nanos(info.modified);
let mut cache = crate::cache::Cache::disabled();
cache.record(&path, info.size, mtime, false, 0xDEAD_BEEF, 0);
let hashed = super::hash_candidates(vec![info], false, 300, false, &cache);
assert_eq!(hashed[0].0, 0xDEAD_BEEF);
}
#[test]
fn hash_candidates_recomputes_when_mtime_differs() {
let root = TempDir::new().unwrap();
let path = root.path().join("f.bin");
let mut f = File::create_new(&path).unwrap();
f.write_all(b"real content for cache miss test").unwrap();
let info = FileInfo::new(path.clone()).unwrap();
let mtime = crate::cache::mtime_nanos(info.modified);
let real = info.initpages_hash(300).unwrap();
let mut cache = crate::cache::Cache::disabled();
cache.record(&path, info.size, mtime + 1, false, 0xDEAD_BEEF, 0);
let hashed = super::hash_candidates(vec![info], false, 300, false, &cache);
assert_eq!(hashed[0].0, real);
assert_ne!(hashed[0].0, 0xDEAD_BEEF);
} }
} }

283
src/resolver.rs Normal file
View File

@@ -0,0 +1,283 @@
use std::cmp::Ordering;
use unicode_segmentation::UnicodeSegmentation;
use anyhow::{bail, Result};
use bytesize::ByteSize;
use crate::fileinfo::FileInfo;
use crate::formatter::Formatter;
use crate::params::Params;
use crate::pipeline::DedupReport;
#[derive(Debug, Clone, Copy, PartialEq, Eq, clap::ValueEnum)]
pub enum KeepStrategy {
Newest,
Oldest,
First,
Last,
Shortest,
Shallowest,
}
fn path_string(file: &FileInfo) -> String {
file.path.to_string_lossy().into_owned()
}
fn path_len(file: &FileInfo) -> usize {
file.path.to_string_lossy().graphemes(true).count()
}
fn depth(file: &FileInfo) -> usize {
file.path.iter().count()
}
fn compare_keep(a: &FileInfo, b: &FileInfo, strategy: KeepStrategy) -> Ordering {
match strategy {
KeepStrategy::Newest => b
.modified
.cmp(&a.modified)
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::Oldest => a
.modified
.cmp(&b.modified)
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::First => path_string(a).cmp(&path_string(b)),
KeepStrategy::Last => path_string(b).cmp(&path_string(a)),
KeepStrategy::Shortest => path_len(a)
.cmp(&path_len(b))
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::Shallowest => depth(a)
.cmp(&depth(b))
.then_with(|| path_len(a).cmp(&path_len(b)))
.then_with(|| path_string(a).cmp(&path_string(b))),
}
}
pub fn select_keeper(files: &[FileInfo], strategy: KeepStrategy) -> usize {
files
.iter()
.enumerate()
.min_by(|(_, a), (_, b)| compare_keep(a, b, strategy))
.map(|(index, _)| index)
.unwrap_or(0)
}
pub fn run(
report: &DedupReport,
strategy: KeepStrategy,
force: bool,
params: &Params,
) -> Result<()> {
if report.groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
}
let group_count = report.groups.len();
let mut victim_count: u64 = 0;
let mut victim_bytes: u64 = 0;
let mut freed_bytes: u64 = 0;
let mut failures: u64 = 0;
for group in &report.groups {
let keeper = select_keeper(&group.files, strategy);
if !force {
println!(
"KEEP {}",
Formatter::human_path(&group.files[keeper], params, report.max_path_len)
.unwrap_or_default()
);
}
for (index, file) in group.files.iter().enumerate() {
if index == keeper {
continue;
}
victim_count += 1;
victim_bytes += file.size;
match force {
false => println!(
"DELETE {} {}",
Formatter::human_path(file, params, report.max_path_len).unwrap_or_default(),
Formatter::human_filesize(file).unwrap_or_default()
),
true => match std::fs::remove_file(&file.path) {
Ok(_) => {
freed_bytes += file.size;
println!("deleted {}", file.path.display());
}
Err(_) => {
failures += 1;
println!("FAILED {}", file.path.display());
}
},
}
}
}
match force {
false => println!(
"\n{group_count} groups, would free {}. Re-run with --force to delete.",
ByteSize::b(victim_bytes)
),
true => println!(
"\ndeleted {} files, freed {} across {group_count} groups.",
victim_count - failures,
ByteSize::b(freed_bytes)
),
}
if failures > 0 {
bail!("{failures} deletion(s) failed");
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use std::path::PathBuf;
use std::time::{Duration, SystemTime};
use crate::params::Params;
use crate::pipeline::{DedupReport, DuplicateGroup};
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
fn file(path: &str, mtime_secs: u64) -> FileInfo {
FileInfo {
path: PathBuf::from(path).into_boxed_path(),
size: 0,
modified: SystemTime::UNIX_EPOCH + Duration::from_secs(mtime_secs),
}
}
#[test]
fn newest_keeps_greatest_mtime() {
let files = vec![file("/a/x", 10), file("/a/y", 30), file("/a/z", 20)];
assert_eq!(select_keeper(&files, KeepStrategy::Newest), 1);
}
#[test]
fn oldest_keeps_least_mtime() {
let files = vec![file("/a/x", 10), file("/a/y", 30), file("/a/z", 20)];
assert_eq!(select_keeper(&files, KeepStrategy::Oldest), 0);
}
#[test]
fn first_keeps_lexicographically_smallest_path() {
let files = vec![file("/a/y", 10), file("/a/x", 10), file("/a/z", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::First), 1);
}
#[test]
fn last_keeps_lexicographically_greatest_path() {
let files = vec![file("/a/y", 10), file("/a/x", 10), file("/a/z", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Last), 2);
}
#[test]
fn shortest_keeps_fewest_chars() {
let files = vec![file("/aaa/bbb", 10), file("/a/b", 10), file("/aa/bb", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Shortest), 1);
}
#[test]
fn shallowest_keeps_fewest_components() {
let files = vec![file("/a/b/c/d", 10), file("/a/b", 10), file("/a/b/c", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Shallowest), 1);
}
#[test]
fn newest_tiebreak_is_smallest_path_regardless_of_order() {
let forward = vec![file("/a/y", 30), file("/a/x", 30)];
let reversed = vec![file("/a/x", 30), file("/a/y", 30)];
assert_eq!(
forward[select_keeper(&forward, KeepStrategy::Newest)].path,
reversed[select_keeper(&reversed, KeepStrategy::Newest)].path
);
assert_eq!(
forward[select_keeper(&forward, KeepStrategy::Newest)]
.path
.to_string_lossy(),
"/a/x"
);
}
fn write_dup_report(root: &TempDir, names: &[&str]) -> DedupReport {
let files = names
.iter()
.map(|name| {
let path = root.path().join(name);
let mut f = File::create_new(&path).unwrap();
f.write_all(b"identical duplicate payload").unwrap();
FileInfo::new(path).unwrap()
})
.collect::<Vec<_>>();
DedupReport {
groups: vec![DuplicateGroup { hash: 0, files }],
max_path_len: 0,
}
}
#[test]
fn force_deletes_victims_and_keeps_the_keeper() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
super::run(&report, KeepStrategy::First, true, &params).unwrap();
assert!(root.path().join("a.bin").exists());
assert!(!root.path().join("b.bin").exists());
}
#[test]
fn dry_run_deletes_nothing() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
super::run(&report, KeepStrategy::First, false, &params).unwrap();
assert!(root.path().join("a.bin").exists());
assert!(root.path().join("b.bin").exists());
}
#[test]
fn force_reports_error_when_a_deletion_fails() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
std::fs::remove_file(root.path().join("b.bin")).unwrap();
let result = super::run(&report, KeepStrategy::First, true, &params);
assert!(result.is_err());
}
#[test]
fn empty_report_is_ok() {
let params = Params::default();
let report = DedupReport {
groups: vec![],
max_path_len: 0,
};
assert!(super::run(&report, KeepStrategy::First, false, &params).is_ok());
}
}

View File

@@ -1,118 +1,48 @@
#![allow(unused)] use crate::{fileinfo::FileInfo, params::Params, pipeline::spinner};
use crate::{fileinfo::FileInfo, params::Params};
use anyhow::Result; use anyhow::Result;
use indicatif::{ProgressBar, ProgressStyle}; use std::path::Path;
use std::{fs, path::PathBuf, time::Duration};
use globwalk::{GlobWalker, GlobWalkerBuilder}; use globwalk::{GlobWalker, GlobWalkerBuilder};
#[derive(Debug, Clone)]
pub struct Scanner { pub struct Scanner {
pub directory: Option<PathBuf>, pub directory: Box<Path>,
pub filetypes: Option<String>,
pub min_depth: Option<usize>, pub min_depth: Option<usize>,
pub max_depth: Option<usize>, pub max_depth: Option<usize>,
pub include_types: Option<String>,
pub exclude_types: Option<String>,
pub min_size: Option<u64>, pub min_size: Option<u64>,
pub follow_links: bool, pub follow_links: bool,
pub progress: bool,
} }
impl Scanner { impl Scanner {
pub fn new() -> Self { pub fn new(app_args: &Params) -> Result<Self> {
Self { Ok(Self {
directory: None, directory: app_args.get_directory()?.into_boxed_path(),
filetypes: None, include_types: app_args.types.clone(),
min_depth: None, exclude_types: app_args.exclude_types.clone(),
max_depth: None, min_depth: app_args.min_depth,
min_size: None, max_depth: app_args.max_depth,
follow_links: true, min_size: app_args.get_min_size(),
} follow_links: app_args.follow_links,
} progress: app_args.progress,
pub fn build(app_args: &Params) -> Result<Self> {
let scan_directory = app_args.get_directory()?;
Ok(Scanner::new())
.map(|scanner| scanner.directory(scan_directory))
.map(|scanner| match app_args.get_min_size() {
Some(min_size) => scanner.min_size(min_size),
None => scanner,
})
.map(|scanner| match app_args.get_types() {
Some(ftypes) => scanner.filetypes(ftypes),
None => scanner,
})
.map(|scanner| match app_args.min_depth {
Some(min_depth) => scanner.min_depth(min_depth),
None => scanner,
})
.map(|scanner| match app_args.max_depth {
Some(max_depth) => scanner.max_depth(max_depth),
None => scanner,
}) })
} }
pub fn min_size(&self, min_size: u64) -> Self { fn scan_patterns(&self) -> Result<Vec<String>> {
Self { let include_types = match &self.include_types {
min_size: Some(min_size), Some(ftypes) => Some(format!("**/*.{{{ftypes}}}")),
..self.clone() None => Some("**/*".to_string()),
}
}
pub fn min_depth(&self, min_depth: usize) -> Self {
Self {
min_depth: Some(min_depth),
..self.clone()
}
}
pub fn max_depth(&self, max_depth: usize) -> Self {
Self {
max_depth: Some(max_depth),
..self.clone()
}
}
pub fn directory(&self, dir: PathBuf) -> Self {
Self {
directory: Some(dir),
..self.clone()
}
}
pub fn filetypes(&self, patterns: String) -> Self {
Self {
filetypes: Some(patterns),
..self.clone()
}
}
pub fn ignore_links(&self) -> Self {
Self {
follow_links: false,
..self.clone()
}
}
pub fn follow_links(&self) -> Self {
Self {
follow_links: true,
..self.clone()
}
}
fn scan_patterns(&self) -> Result<String> {
Ok(match self.filetypes.clone() {
Some(ftypes) => format!("**/*{{{ftypes}}}"),
None => "**/*".to_string(),
})
}
fn scan_dir(&self) -> Result<PathBuf> {
let scan_dir = match self.directory.clone() {
Some(path) => path,
None => std::env::current_dir()?,
}; };
Ok(fs::canonicalize(scan_dir)?) let exclude_types = self
.exclude_types
.as_ref()
.map(|ftypes| format!("!**/*.{{{ftypes}}}"));
Ok(vec![include_types, exclude_types]
.into_iter()
.flatten()
.collect())
} }
fn attach_link_opts(&self, walker: GlobWalkerBuilder) -> Result<GlobWalkerBuilder> { fn attach_link_opts(&self, walker: GlobWalkerBuilder) -> Result<GlobWalkerBuilder> {
@@ -132,10 +62,11 @@ impl Scanner {
None => Ok(walker), None => Ok(walker),
} }
} }
fn build_walker(&self) -> Result<GlobWalker> { fn build_walker(&self) -> Result<GlobWalker> {
let walker = Ok(GlobWalkerBuilder::from_patterns( let walker = Ok(GlobWalkerBuilder::from_patterns(
self.scan_dir()?, self.directory.clone(),
&[self.scan_patterns()?], &self.scan_patterns()?,
)) ))
.and_then(|walker| self.attach_walker_min_depth(walker)) .and_then(|walker| self.attach_walker_min_depth(walker))
.and_then(|walker| self.attach_walker_max_depth(walker)) .and_then(|walker| self.attach_walker_max_depth(walker))
@@ -145,29 +76,141 @@ impl Scanner {
} }
pub fn scan(&self) -> Result<Vec<FileInfo>> { pub fn scan(&self) -> Result<Vec<FileInfo>> {
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")?; let bar = spinner(self.progress, "paths mapped");
let progress_bar = ProgressBar::new_spinner(); let min_size = self.min_size.unwrap_or(0);
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("paths mapped");
let min_size = self.min_size.unwrap_or_default();
let results = self let files: Vec<FileInfo> = self
.build_walker()? .build_walker()?
.filter_map(Result::ok) .filter_map(Result::ok)
.map(|entity| entity.into_path()) .map(|entity| entity.into_path())
.map(|path| { .inspect(|_path| bar.inc(1))
progress_bar.inc(1);
path
})
.filter(|path| path.is_file()) .filter(|path| path.is_file())
.map(FileInfo::new) .map(FileInfo::new)
.filter_map(Result::ok) .filter_map(Result::ok)
.filter(|file| file.size > min_size) .filter(|file| file.size >= min_size)
.collect::<Vec<FileInfo>>(); .collect();
progress_bar.finish_with_message("paths mapped"); bar.finish_with_message("paths mapped");
Ok(files)
Ok(results) }
}
#[cfg(test)]
mod tests {
use crate::params::Params;
use std::fs::File;
use tempfile::TempDir;
use super::Scanner;
#[test]
fn ensure_file_include_type_filter_includes_expected_file_types() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
types: Some(String::from("js,csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-css-file.css").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
}
#[test]
fn ensure_file_exclude_type_filter_excludes_expected_file_types() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
exclude_types: Some(String::from("js,csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-css-file.css").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
}
#[test]
fn complex_file_type_params() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
types: Some(String::from("js,csv,rs")),
exclude_types: Some(String::from("csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
} }
} }

View File

@@ -1,44 +0,0 @@
use anyhow::Result;
use std::fs::File;
use std::io::Read;
use std::os::unix::fs::MetadataExt;
use std::sync::Arc;
use uuid::Uuid;
const PARTIAL_SIZE: u64 = 4096;
#[allow(unused)]
#[derive(Clone, Debug)]
pub struct FileMeta {
pub id: Uuid,
pub path: Box<str>,
pub size: u64,
pub modtime: i64,
pub partial: Arc<[u8]>,
pub full_hash: Arc<[u8]>,
}
impl FileMeta {
fn create_partial(path: &str) -> Result<Arc<[u8]>> {
let mut partial_take = File::open(path)?.take(PARTIAL_SIZE);
let mut partial_buffer = Vec::with_capacity(PARTIAL_SIZE as usize);
partial_take.read_to_end(&mut partial_buffer)?;
Ok(Arc::from(partial_buffer))
}
pub fn new(path: Box<str>) -> Result<Self> {
let pstr = path.to_string();
let filemeta = std::fs::metadata(&pstr)?;
let partial = Self::create_partial(&pstr)?;
Ok(Self {
id: Uuid::new_v4(),
size: filemeta.size(),
modtime: filemeta.mtime(),
full_hash: Arc::new([]),
path,
partial,
})
}
}

View File

@@ -1,79 +0,0 @@
pub mod file;
mod processor;
mod scanner;
mod store;
use anyhow::Result;
use std::sync::mpsc;
use std::sync::mpsc::channel;
use std::sync::Arc;
use std::sync::Mutex;
use threadpool::ThreadPool;
use self::processor::Processor;
use self::scanner::Scanner;
use self::store::Store;
pub type FileQueue = Arc<Mutex<Vec<Box<str>>>>;
pub enum Message {
AddScanDirectory(Box<str>),
Exit,
None,
}
pub struct Server {
pub fq: FileQueue,
pub dupstore: Arc<Store>,
pub tpool: ThreadPool,
}
impl Server {
pub fn new() -> Result<Self> {
Ok(Self {
fq: Arc::new(Mutex::new(vec![])),
dupstore: Arc::new(Store::new()),
tpool: ThreadPool::new(4),
})
}
pub fn start(&self, rx: mpsc::Receiver<Message>) -> Result<()> {
let processor_fq = self.fq.clone();
let processor_store = self.dupstore.clone();
let (processor_tx, processor_rx) = channel::<Message>();
self.tpool.execute(move || {
Processor::new(processor_fq, processor_store, processor_rx)
.process()
.expect("processer execution interrupted.");
});
let scanner_fq = self.fq.clone();
let (scanner_tx, scanner_rx) = channel::<Message>();
self.tpool.execute(move || {
Scanner::new(scanner_fq, scanner_rx)
.index()
.expect("scanner indexing interrupted.");
});
self.tpool.execute(move || loop {
match rx.recv() {
Ok(Message::AddScanDirectory(path)) => {
scanner_tx
.send(Message::AddScanDirectory(path))
.expect("scanner tx message passing failed.");
}
Ok(Message::None) => {}
Ok(Message::Exit) | Err(_) => {
scanner_tx.send(Message::Exit).unwrap_or_default();
processor_tx.send(Message::Exit).unwrap_or_default();
break;
}
};
});
self.tpool.join();
Ok(())
}
}

View File

@@ -1,146 +0,0 @@
use anyhow::Result;
use std::sync::mpsc::{Receiver, TryRecvError};
use std::sync::Arc;
use super::file::FileMeta;
use super::store::{Index, Store};
use super::{FileQueue, Message};
pub struct Processor {
files: FileQueue,
duplicates: Arc<Store>,
msg_rx: Receiver<Message>,
}
impl Processor {
pub fn new(files: FileQueue, duplicates: Arc<Store>, msg_rx: Receiver<Message>) -> Self {
Self {
files,
duplicates,
msg_rx,
}
}
pub fn process(&self) -> Result<()> {
loop {
match self.msg_rx.try_recv() {
Ok(Message::Exit) => break,
Err(TryRecvError::Empty) => {},
_ => {}
}
let next_file = {
let mut cfiles = self.files.lock().unwrap();
cfiles.pop()
};
match next_file {
None => continue,
Some(file_res) => match FileMeta::new(file_res) {
Err(_) => continue,
Ok(fm) => {
let fm_arc = Arc::new(fm);
self.duplicates
.add(Index::Size(fm_arc.size), fm_arc.clone());
}
},
}
}
Ok(())
}
}
#[cfg(test)]
mod tests {
use super::*;
use anyhow::Result;
use std::sync::mpsc;
use std::sync::{Arc, Mutex};
use std::thread;
use tempfile::TempDir;
use std::fs::File;
use std::io::Write;
#[test]
fn processor_differentiates_files_with_different_sizes() -> Result<()> {
let file_queue = Arc::new(Mutex::new(vec![]));
let (tx, rx) = mpsc::channel::<Message>();
let root = TempDir::new()?;
let store = Arc::new(Store::new());
let fq_c = file_queue.clone();
let store_c = store.clone();
let proc_thread = thread::spawn(move || {
let processor = Processor::new(fq_c, store_c, rx);
processor.process().expect("processor failed");
});
let files =
[("hello.txt", "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat"),
("hello_dup.txt", "Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum")];
for (filename, content) in files.into_iter() {
let fpath = root.path().join(filename);
let mut tf = File::create(fpath.clone())?;
tf.write_all(content.as_bytes())?;
let mut mfq = file_queue.lock().unwrap();
mfq.push(fpath.into_os_string().into_string().unwrap().into_boxed_str());
}
for _ in 0..10 {
if store.entries().len() < 2 {
thread::sleep(std::time::Duration::from_millis(100));
}
}
tx.send(Message::Exit).expect("unable to send msg to processor");
proc_thread.join().expect("failed to join on thread!");
assert!(store.entries().len() == 2);
Ok(())
}
#[test]
fn processor_groups_files_with_same_size() -> Result<()> {
let file_queue = Arc::new(Mutex::new(vec![]));
let (tx, rx) = mpsc::channel::<Message>();
let root = TempDir::new()?;
let store = Arc::new(Store::new());
let fq_c = file_queue.clone();
let store_c = store.clone();
let proc_thread = thread::spawn(move || {
let processor = Processor::new(fq_c, store_c, rx);
processor.process().expect("processor failed");
});
let files =
[("hello.txt", "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum"),
("hello_dup.txt", "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum")];
for (filename, content) in files.into_iter() {
let fpath = root.path().join(filename);
let mut tf = File::create(fpath.clone())?;
tf.write_all(content.as_bytes())?;
let mut mfq = file_queue.lock().unwrap();
mfq.push(fpath.into_os_string().into_string().unwrap().into_boxed_str());
}
for _ in 0..10 {
if store.entries().is_empty() {
thread::sleep(std::time::Duration::from_millis(100));
}
}
tx.send(Message::Exit).expect("unable to send msg to processor");
proc_thread.join().expect("failed to join on thread!");
assert!(store.entries().len() == 1);
Ok(())
}
}

View File

@@ -1,128 +0,0 @@
use super::{FileQueue, Message};
use anyhow::Result;
use std::fs;
use std::sync::mpsc::{self, Receiver};
use std::sync::Arc;
use std::sync::Mutex;
#[derive(Debug)]
pub struct Scanner {
files: FileQueue,
proc_queue: FileQueue,
msg_rx: Receiver<Message>,
}
impl Scanner {
pub fn new(fq: FileQueue, rx: Receiver<Message>) -> Self {
Self {
files: fq,
proc_queue: Arc::new(Mutex::new(vec![])),
msg_rx: rx,
}
}
pub fn index(&self) -> Result<()> {
loop {
match self.msg_rx.try_recv() {
Ok(Message::AddScanDirectory(path)) => {
let mut mfq = self.proc_queue.lock().unwrap();
mfq.push(path);
}
Err(mpsc::TryRecvError::Empty) | Ok(Message::None) => {}
Ok(Message::Exit) | Err(_) => break,
}
let npath = {
match self.proc_queue.try_lock() {
Ok(mut q) => q.pop(),
Err(_) => None,
}
};
match npath {
None => continue,
Some(path) => {
std::fs::read_dir(path.as_ref())?
.filter_map(Result::ok)
.for_each(|entry: fs::DirEntry| {
let mdata =
fs::metadata(entry.path()).expect("unable to read file metadata.");
let mpath = entry
.path()
.into_os_string()
.into_string()
.expect("invalid path conversion failed.")
.into_boxed_str();
match mdata.is_dir() {
true => {
let mut pq = self
.proc_queue
.lock()
.expect("proc queue lock acq failed.");
pq.push(mpath);
}
false => {
let mut fq = self
.files
.lock()
.expect("file queue lock acq failed.");
fq.push(mpath)
}
}
});
}
}
}
Ok(())
}
}
#[cfg(test)]
mod tests {
use super::*;
use anyhow::Result;
use std::fs::File;
use std::io::Write;
use std::sync::mpsc::channel;
use tempfile::TempDir;
use std::thread;
#[test]
fn scanner_scans_files() -> Result<()> {
let files =
[("hello.txt", "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum"),
("hello_dup.txt", "Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum")];
let file_queue: FileQueue = Arc::new(Mutex::new(vec![]));
let (tx, rx) = channel::<Message>();
let root = TempDir::new()?;
for (filename, content) in files.into_iter() {
let mut tf = File::create(root.path().join(filename))?;
tf.write_all(content.as_bytes())?;
}
let rpath = root.path().to_str().unwrap().to_string().into_boxed_str();
let fq_clone = file_queue.clone();
let scanner_thread = thread::spawn(move || {
let scanner = Scanner::new(fq_clone, rx);
scanner.index().unwrap();
});
tx.send(Message::AddScanDirectory(rpath))?;
tx.send(Message::Exit)?;
scanner_thread.join().unwrap();
let fq_len = {
let v = file_queue.lock().unwrap();
v.len()
};
assert_eq!(files.len(), fq_len);
Ok(())
}
}

View File

@@ -1,35 +0,0 @@
use super::file::FileMeta;
use dashmap::DashMap;
use std::sync::Arc;
#[allow(unused)]
#[derive(Debug, Hash, PartialEq, Eq)]
pub enum Index {
Size(u64),
Partial(Arc<[u8]>),
Full(Box<str>),
}
#[derive(Debug)]
pub struct Store {
internal: Arc<DashMap<Index, Vec<Arc<FileMeta>>>>,
}
impl Store {
pub fn new() -> Self {
Self {
internal: Arc::new(DashMap::new()),
}
}
pub fn entries(&self) -> Arc<DashMap<Index, Vec<Arc<FileMeta>>>> {
self.internal.clone()
}
pub fn add(&self, index: Index, file: Arc<FileMeta>) {
self.internal
.entry(index)
.and_modify(|fg| fg.push(file.clone()))
.or_insert(vec![file]);
}
}

499
src/tui/app.rs Normal file
View File

@@ -0,0 +1,499 @@
use std::path::PathBuf;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use crate::resolver::{select_keeper, KeepStrategy};
pub const STRATEGIES: [KeepStrategy; 6] = [
KeepStrategy::Newest,
KeepStrategy::Oldest,
KeepStrategy::First,
KeepStrategy::Last,
KeepStrategy::Shortest,
KeepStrategy::Shallowest,
];
pub fn strategy_label(strategy: KeepStrategy) -> &'static str {
match strategy {
KeepStrategy::Newest => "newest",
KeepStrategy::Oldest => "oldest",
KeepStrategy::First => "first",
KeepStrategy::Last => "last",
KeepStrategy::Shortest => "shortest",
KeepStrategy::Shallowest => "shallowest",
}
}
#[derive(Clone, Copy, PartialEq)]
pub enum Focus {
Groups,
Files,
}
#[derive(Clone, Copy, PartialEq)]
pub enum StrategyScope {
CurrentGroup,
AllGroups,
}
#[derive(Clone, Copy, PartialEq)]
pub enum Popup {
None,
Strategy { scope: StrategyScope },
ConfirmDelete,
}
pub enum Key {
Up,
Down,
Tab,
Space,
Enter,
Esc,
Char(char),
}
pub enum Outcome {
Continue,
Quit,
Delete,
Open(PathBuf),
}
pub struct Group {
pub files: Vec<FileInfo>,
pub marked: Vec<bool>,
}
impl Group {
pub fn size(&self) -> u64 {
self.files.first().map(|f| f.size).unwrap_or(0)
}
}
pub struct App {
pub groups: Vec<Group>,
pub group_cursor: usize,
pub file_cursor: usize,
pub focus: Focus,
pub popup: Popup,
pub strategy_cursor: usize,
pub status: Option<String>,
pub should_quit: bool,
}
fn clamp_add(cur: usize, delta: isize, max: usize) -> usize {
(cur as isize + delta).clamp(0, max as isize) as usize
}
impl App {
pub fn from_groups(groups: Vec<DuplicateGroup>) -> Self {
let groups = groups
.into_iter()
.map(|g| Group {
marked: vec![false; g.files.len()],
files: g.files,
})
.collect();
Self {
groups,
group_cursor: 0,
file_cursor: 0,
focus: Focus::Files,
popup: Popup::None,
strategy_cursor: 0,
status: None,
should_quit: false,
}
}
pub fn move_group(&mut self, delta: isize) {
if self.groups.is_empty() {
return;
}
self.group_cursor = clamp_add(self.group_cursor, delta, self.groups.len() - 1);
let len = self.groups[self.group_cursor].files.len();
self.file_cursor = self.file_cursor.min(len.saturating_sub(1));
}
pub fn move_file(&mut self, delta: isize) {
if let Some(g) = self.groups.get(self.group_cursor) {
self.file_cursor = clamp_add(self.file_cursor, delta, g.files.len().saturating_sub(1));
}
}
pub fn toggle_mark(&mut self) {
let fc = self.file_cursor;
if let Some(g) = self.groups.get_mut(self.group_cursor) {
if fc >= g.marked.len() {
return;
}
if !g.marked[fc] && g.marked.iter().filter(|m| !**m).count() <= 1 {
return;
}
g.marked[fc] = !g.marked[fc];
}
}
pub fn apply_strategy(&mut self, strategy: KeepStrategy, scope: StrategyScope) {
let indices: Vec<usize> = match scope {
StrategyScope::CurrentGroup => vec![self.group_cursor],
StrategyScope::AllGroups => (0..self.groups.len()).collect(),
};
for gi in indices {
if let Some(g) = self.groups.get_mut(gi) {
if g.files.len() < 2 {
continue;
}
let keeper = select_keeper(&g.files, strategy);
for i in 0..g.marked.len() {
g.marked[i] = i != keeper;
}
}
}
}
pub fn marked_count(&self) -> usize {
self.groups
.iter()
.flat_map(|g| g.marked.iter())
.filter(|m| **m)
.count()
}
pub fn reclaimable_bytes(&self) -> u64 {
self.groups
.iter()
.flat_map(|g| g.files.iter().zip(&g.marked))
.filter(|(_, m)| **m)
.map(|(f, _)| f.size)
.sum()
}
pub fn marked_paths(&self) -> Vec<PathBuf> {
self.groups
.iter()
.flat_map(|g| g.files.iter().zip(&g.marked))
.filter(|(_, m)| **m)
.map(|(f, _)| f.path.to_path_buf())
.collect()
}
pub fn apply_deletion(&mut self, deleted: &[PathBuf]) {
use std::collections::HashSet;
let removed: HashSet<&std::path::Path> = deleted.iter().map(|p| p.as_path()).collect();
for g in &mut self.groups {
let mut files = Vec::new();
let mut marked = Vec::new();
for (i, f) in g.files.iter().enumerate() {
if !removed.contains(&*f.path) {
files.push(f.clone());
marked.push(g.marked[i]);
}
}
g.files = files;
g.marked = marked;
}
self.groups.retain(|g| g.files.len() >= 2);
if self.groups.is_empty() {
self.should_quit = true;
self.group_cursor = 0;
self.file_cursor = 0;
return;
}
self.group_cursor = self.group_cursor.min(self.groups.len() - 1);
let len = self.groups[self.group_cursor].files.len();
self.file_cursor = self.file_cursor.min(len.saturating_sub(1));
}
pub fn handle_key(&mut self, key: Key) -> Outcome {
match self.popup {
Popup::Strategy { scope } => self.handle_strategy_key(key, scope),
Popup::ConfirmDelete => self.handle_confirm_key(key),
Popup::None => self.handle_main_key(key),
}
}
fn handle_main_key(&mut self, key: Key) -> Outcome {
self.status = None;
match key {
Key::Char('q') | Key::Esc => {
self.should_quit = true;
Outcome::Quit
}
Key::Tab => {
self.focus = match self.focus {
Focus::Groups => Focus::Files,
Focus::Files => Focus::Groups,
};
Outcome::Continue
}
Key::Up | Key::Char('k') => {
self.move_in_focus(-1);
Outcome::Continue
}
Key::Down | Key::Char('j') => {
self.move_in_focus(1);
Outcome::Continue
}
Key::Space => {
if self.focus == Focus::Files {
self.toggle_mark();
}
Outcome::Continue
}
Key::Char('s') => {
self.open_strategy(StrategyScope::CurrentGroup);
Outcome::Continue
}
Key::Char('S') => {
self.open_strategy(StrategyScope::AllGroups);
Outcome::Continue
}
Key::Char('o') => self.open_current(),
Key::Char('d') => {
if self.marked_count() > 0 {
self.popup = Popup::ConfirmDelete;
} else {
self.status = Some("nothing marked".to_string());
}
Outcome::Continue
}
_ => Outcome::Continue,
}
}
fn handle_strategy_key(&mut self, key: Key, scope: StrategyScope) -> Outcome {
match key {
Key::Up | Key::Char('k') => {
self.strategy_cursor = self.strategy_cursor.saturating_sub(1);
}
Key::Down | Key::Char('j') => {
self.strategy_cursor = (self.strategy_cursor + 1).min(STRATEGIES.len() - 1);
}
Key::Enter => {
let strategy = STRATEGIES[self.strategy_cursor];
self.apply_strategy(strategy, scope);
self.popup = Popup::None;
}
Key::Esc => self.popup = Popup::None,
_ => {}
}
Outcome::Continue
}
fn handle_confirm_key(&mut self, key: Key) -> Outcome {
match key {
Key::Char('y') => {
self.popup = Popup::None;
Outcome::Delete
}
Key::Char('n') | Key::Esc => {
self.popup = Popup::None;
Outcome::Continue
}
_ => Outcome::Continue,
}
}
fn move_in_focus(&mut self, delta: isize) {
match self.focus {
Focus::Groups => self.move_group(delta),
Focus::Files => self.move_file(delta),
}
}
fn open_strategy(&mut self, scope: StrategyScope) {
self.strategy_cursor = 0;
self.popup = Popup::Strategy { scope };
}
fn open_current(&self) -> Outcome {
match self.groups.get(self.group_cursor).and_then(|g| g.files.get(self.file_cursor)) {
Some(f) => Outcome::Open(f.path.to_path_buf()),
None => Outcome::Continue,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use std::path::PathBuf;
use std::time::{Duration, UNIX_EPOCH};
fn file(path: &str, size: u64, mtime_secs: u64) -> FileInfo {
FileInfo {
path: PathBuf::from(path).into_boxed_path(),
size,
modified: UNIX_EPOCH + Duration::from_secs(mtime_secs),
}
}
fn group(files: Vec<FileInfo>) -> DuplicateGroup {
DuplicateGroup { hash: 0, files }
}
fn sample() -> App {
App::from_groups(vec![
group(vec![file("/a/x", 10, 100), file("/a/y", 10, 200)]),
group(vec![file("/b/p", 5, 50), file("/b/q", 5, 60), file("/b/r", 5, 70)]),
])
}
#[test]
fn move_group_clamps_at_bounds() {
let mut app = sample();
app.move_group(-1);
assert_eq!(app.group_cursor, 0);
app.move_group(1);
assert_eq!(app.group_cursor, 1);
app.move_group(5);
assert_eq!(app.group_cursor, 1);
}
#[test]
fn move_group_reclamps_file_cursor() {
let mut app = sample();
app.group_cursor = 1;
app.file_cursor = 2;
app.move_group(-1);
assert_eq!(app.group_cursor, 0);
assert_eq!(app.file_cursor, 1);
}
#[test]
fn move_file_clamps() {
let mut app = sample();
app.move_file(-1);
assert_eq!(app.file_cursor, 0);
app.move_file(10);
assert_eq!(app.file_cursor, 1);
}
#[test]
fn toggle_mark_blocks_marking_the_last_survivor() {
let mut app = sample();
app.group_cursor = 0;
app.file_cursor = 0;
app.toggle_mark();
assert!(app.groups[0].marked[0]);
app.file_cursor = 1;
app.toggle_mark();
assert!(!app.groups[0].marked[1]);
}
#[test]
fn toggle_mark_allows_unmark() {
let mut app = sample();
app.file_cursor = 0;
app.toggle_mark();
app.toggle_mark();
assert!(!app.groups[0].marked[0]);
}
#[test]
fn apply_strategy_current_group_marks_all_but_keeper() {
let mut app = sample();
app.group_cursor = 0;
app.apply_strategy(KeepStrategy::Newest, StrategyScope::CurrentGroup);
assert_eq!(app.groups[0].marked, vec![true, false]);
assert_eq!(app.groups[1].marked, vec![false, false, false]);
}
#[test]
fn apply_strategy_all_groups() {
let mut app = sample();
app.apply_strategy(KeepStrategy::Oldest, StrategyScope::AllGroups);
assert_eq!(app.groups[0].marked, vec![false, true]);
assert_eq!(app.groups[1].marked, vec![false, true, true]);
}
#[test]
fn totals_sum_marked_only() {
let mut app = sample();
app.apply_strategy(KeepStrategy::Newest, StrategyScope::AllGroups);
assert_eq!(app.marked_count(), 3);
assert_eq!(app.reclaimable_bytes(), 20);
}
#[test]
fn marked_paths_lists_marked_files() {
let mut app = sample();
app.group_cursor = 0;
app.file_cursor = 0;
app.toggle_mark();
assert_eq!(app.marked_paths(), vec![PathBuf::from("/a/x")]);
}
#[test]
fn apply_deletion_removes_files_and_drops_small_groups() {
let mut app = sample();
app.apply_deletion(&[PathBuf::from("/a/x")]);
assert_eq!(app.groups.len(), 1);
assert_eq!(app.groups[0].files.len(), 3);
}
#[test]
fn apply_deletion_emptying_everything_sets_quit() {
let mut app = App::from_groups(vec![group(vec![file("/a/x", 1, 1), file("/a/y", 1, 2)])]);
app.apply_deletion(&[PathBuf::from("/a/x")]);
assert!(app.groups.is_empty());
assert!(app.should_quit);
}
#[test]
fn handle_key_quits_on_q() {
let mut app = sample();
assert!(matches!(app.handle_key(Key::Char('q')), Outcome::Quit));
assert!(app.should_quit);
}
#[test]
fn handle_key_space_marks_in_files_focus() {
let mut app = sample();
app.focus = Focus::Files;
app.handle_key(Key::Space);
assert!(app.groups[0].marked[0]);
}
#[test]
fn handle_key_strategy_popup_apply_flow() {
let mut app = sample();
app.handle_key(Key::Char('S'));
assert!(matches!(app.popup, Popup::Strategy { .. }));
app.handle_key(Key::Enter);
assert!(matches!(app.popup, Popup::None));
assert_eq!(app.groups[0].marked, vec![true, false]);
}
#[test]
fn handle_key_delete_requires_marks_then_confirms() {
let mut app = sample();
app.handle_key(Key::Char('d'));
assert!(matches!(app.popup, Popup::None));
app.focus = Focus::Files;
app.handle_key(Key::Space);
app.handle_key(Key::Char('d'));
assert!(matches!(app.popup, Popup::ConfirmDelete));
assert!(matches!(app.handle_key(Key::Char('y')), Outcome::Delete));
assert!(matches!(app.popup, Popup::None));
}
#[test]
fn handle_key_open_returns_path() {
let mut app = sample();
app.focus = Focus::Files;
match app.handle_key(Key::Char('o')) {
Outcome::Open(p) => assert_eq!(p, PathBuf::from("/a/x")),
_ => panic!("expected Open"),
}
}
}

View File

@@ -1,99 +1,129 @@
use std::sync::mpsc::Sender; pub mod app;
use std::time::Duration; mod ui;
use anyhow::{Context, Result}; use std::io::{self, Stdout};
use ratatui::crossterm::event::{self, Event, KeyCode};
use ratatui::layout::{Constraint, Direction, Layout};
use ratatui::widgets::{Block, Borders, List, ListItem, Paragraph};
use ratatui::{DefaultTerminal, Frame};
use crate::server::{Message, Server}; use anyhow::Result;
use std::sync::Arc; use ratatui::backend::CrosstermBackend;
use ratatui::crossterm::event::{self, Event, KeyCode, KeyEvent, KeyEventKind};
use ratatui::crossterm::execute;
use ratatui::crossterm::terminal::{
disable_raw_mode, enable_raw_mode, EnterAlternateScreen, LeaveAlternateScreen,
};
use ratatui::Terminal;
pub struct Tui { use crate::params::Params;
app_tx: Sender<Message>, use crate::pipeline::DedupReport;
server: Arc<Server>, use app::{App, Key, Outcome};
type Tui = Terminal<CrosstermBackend<Stdout>>;
pub fn run(report: DedupReport, params: &Params) -> Result<()> {
if report.groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
} }
impl Tui { let mut app = App::from_groups(report.groups);
pub fn new(app_tx: Sender<Message>, server: Arc<Server>) -> Self { let mut terminal = setup_terminal()?;
Self { app_tx, server } let result = event_loop(&mut terminal, &mut app, params);
restore_terminal(&mut terminal)?;
result
} }
pub fn start(&mut self) -> Result<()> { fn setup_terminal() -> Result<Tui> {
let terminal = ratatui::init(); install_panic_hook();
self.run(terminal).expect("ui loop failed."); enable_raw_mode()?;
ratatui::restore(); let mut stdout = io::stdout();
execute!(stdout, EnterAlternateScreen)?;
Ok(Terminal::new(CrosstermBackend::new(stdout))?)
}
fn restore_terminal(terminal: &mut Tui) -> Result<()> {
let _ = disable_raw_mode();
let _ = execute!(terminal.backend_mut(), LeaveAlternateScreen);
let _ = terminal.show_cursor();
Ok(()) Ok(())
} }
fn poll_events() -> Result<Message> { fn install_panic_hook() {
match event::poll(Duration::from_millis(100)).context("event polling failed.")? { let original = std::panic::take_hook();
true => match event::read().context("event read failed.")? { std::panic::set_hook(Box::new(move |info| {
Event::Key(key) => match key.code { let _ = disable_raw_mode();
KeyCode::Char('q') => Ok(Message::Exit), let _ = execute!(io::stdout(), LeaveAlternateScreen);
_ => Ok(Message::None), original(info);
}, }));
_ => Ok(Message::None),
},
false => Ok(Message::None),
}
} }
fn handle_events(&self) -> Result<Message> { fn event_loop(terminal: &mut Tui, app: &mut App, params: &Params) -> Result<()> {
match Self::poll_events() {
Ok(Message::Exit) => {
self.app_tx
.send(Message::Exit)
.expect("app event send failed.");
Ok(Message::Exit)
}
_ => Ok(Message::None),
}
}
fn draw(&mut self, frame: &mut Frame) {
let listitems = self
.server
.dupstore
.entries()
.iter()
.flat_map(|val| {
let mut group = vec![ListItem::new(format!("{:?}", val.key()))];
val.value().iter().for_each(|f| {
group.push(ListItem::new(format!("|-{}", f.path)));
});
group
})
.collect::<Vec<ListItem>>();
let list = List::new(listitems.clone());
let layout_chunks = Layout::default()
.direction(Direction::Horizontal)
.constraints([Constraint::Percentage(50), Constraint::Percentage(50)])
.split(frame.area());
frame.render_widget(
list.block(Block::new().borders(Borders::ALL)),
layout_chunks[0],
);
frame.render_widget(
Paragraph::new(String::new()).block(Block::new().borders(Borders::ALL)),
layout_chunks[1],
);
}
fn run(&mut self, mut terminal: DefaultTerminal) -> Result<()> {
loop { loop {
terminal.draw(|f| self.draw(f))?; terminal.draw(|frame| ui::draw(frame, app, params))?;
match self.handle_events() { if app.should_quit {
Ok(Message::Exit) => break, break;
_ => {} }
if let Event::Key(key) = event::read()? {
if key.kind != KeyEventKind::Press {
continue;
}
if let Some(translated) = translate(key) {
match app.handle_key(translated) {
Outcome::Quit => {}
Outcome::Continue => {}
Outcome::Delete => perform_deletion(app),
Outcome::Open(path) => {
if let Err(err) = open::that(&path) {
app.status = Some(format!("open failed: {err}"));
}
}
}
} }
} }
if app.should_quit {
break;
}
}
Ok(()) Ok(())
} }
fn translate(key: KeyEvent) -> Option<Key> {
Some(match key.code {
KeyCode::Up => Key::Up,
KeyCode::Down => Key::Down,
KeyCode::Tab => Key::Tab,
KeyCode::Enter => Key::Enter,
KeyCode::Esc => Key::Esc,
KeyCode::Char(' ') => Key::Space,
KeyCode::Char(c) => Key::Char(c),
_ => return None,
})
}
fn perform_deletion(app: &mut App) {
let paths = app.marked_paths();
let mut deleted = Vec::new();
let mut freed: u64 = 0;
let mut failures: u64 = 0;
for path in &paths {
let size = std::fs::metadata(path).map(|m| m.len()).unwrap_or(0);
match std::fs::remove_file(path) {
Ok(_) => {
deleted.push(path.clone());
freed += size;
}
Err(_) => failures += 1,
}
}
app.apply_deletion(&deleted);
app.status = Some(match failures {
0 => format!("deleted {}, freed {}", deleted.len(), bytesize::ByteSize::b(freed)),
_ => format!(
"deleted {}, {failures} failed, freed {}",
deleted.len(),
bytesize::ByteSize::b(freed)
),
});
} }

211
src/tui/ui.rs Normal file
View File

@@ -0,0 +1,211 @@
use ratatui::layout::{Constraint, Direction, Layout, Rect};
use ratatui::style::Style;
use ratatui::text::Line;
use ratatui::widgets::{Block, Borders, Clear, List, ListItem, ListState, Paragraph};
use ratatui::Frame;
use crate::params::Params;
use crate::formatter::Formatter;
use super::app::{strategy_label, App, Focus, Popup, StrategyScope, STRATEGIES};
pub fn draw(frame: &mut Frame, app: &App, params: &Params) {
let rows = Layout::default()
.direction(Direction::Vertical)
.constraints([Constraint::Min(1), Constraint::Length(1)])
.split(frame.area());
let panes = Layout::default()
.direction(Direction::Horizontal)
.constraints([Constraint::Percentage(35), Constraint::Percentage(65)])
.split(rows[0]);
draw_groups(frame, app, panes[0]);
draw_files(frame, app, params, panes[1]);
draw_footer(frame, app, rows[1]);
match app.popup {
Popup::Strategy { scope } => draw_strategy_popup(frame, app, scope),
Popup::ConfirmDelete => draw_confirm_popup(frame, app),
Popup::None => {}
}
}
fn pane_block(title: &str, focused: bool) -> Block<'_> {
let block = Block::default().borders(Borders::ALL).title(title.to_string());
match focused {
true => block.border_style(Style::new().yellow()),
false => block,
}
}
fn draw_groups(frame: &mut Frame, app: &App, area: Rect) {
let items: Vec<ListItem> = app
.groups
.iter()
.enumerate()
.map(|(i, g)| {
ListItem::new(format!(
"{:>3} {} files {}",
i + 1,
g.files.len(),
bytesize::ByteSize::b(g.size())
))
})
.collect();
let list = List::new(items)
.block(pane_block("Groups", app.focus == Focus::Groups))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.group_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_files(frame: &mut Frame, app: &App, params: &Params, area: Rect) {
let title = format!(
"Group {}/{}",
(app.group_cursor + 1).min(app.groups.len().max(1)),
app.groups.len()
);
let items: Vec<ListItem> = match app.groups.get(app.group_cursor) {
Some(g) => g
.files
.iter()
.enumerate()
.map(|(i, f)| {
let box_char = if g.marked[i] { "[x]" } else { "[ ]" };
let path = Formatter::human_path(f, params, 0).unwrap_or_default();
let size = Formatter::human_filesize(f).unwrap_or_default();
let mtime = Formatter::human_mtime(f).unwrap_or_default();
let line = format!("{box_char} {} {} {}", path.trim_end(), size.trim_start(), mtime);
match g.marked[i] {
true => ListItem::new(Line::from(line).style(Style::new().red())),
false => ListItem::new(line),
}
})
.collect(),
None => Vec::new(),
};
let list = List::new(items)
.block(pane_block(&title, app.focus == Focus::Files))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.file_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_footer(frame: &mut Frame, app: &App, area: Rect) {
let hints = "[Tab]pane [Space]mark [s]group [S]all [o]pen [d]elete [q]uit";
let left = match &app.status {
Some(s) => s.clone(),
None => format!(
"marked {} · reclaim {}",
app.marked_count(),
bytesize::ByteSize::b(app.reclaimable_bytes())
),
};
let text = format!("{left} {hints}");
frame.render_widget(Paragraph::new(text), area);
}
fn centered_rect(width: u16, height: u16, area: Rect) -> Rect {
let x = area.x + area.width.saturating_sub(width) / 2;
let y = area.y + area.height.saturating_sub(height) / 2;
Rect {
x,
y,
width: width.min(area.width),
height: height.min(area.height),
}
}
fn draw_strategy_popup(frame: &mut Frame, app: &App, scope: StrategyScope) {
let title = match scope {
StrategyScope::CurrentGroup => "Keep in this group",
StrategyScope::AllGroups => "Keep in ALL groups",
};
let items: Vec<ListItem> = STRATEGIES
.iter()
.map(|s| ListItem::new(strategy_label(*s)))
.collect();
let area = centered_rect(28, (STRATEGIES.len() as u16) + 2, frame.area());
frame.render_widget(Clear, area);
let list = List::new(items)
.block(Block::default().borders(Borders::ALL).title(title))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.strategy_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_confirm_popup(frame: &mut Frame, app: &App) {
let area = centered_rect(48, 3, frame.area());
frame.render_widget(Clear, area);
let text = format!(
"Delete {} files, reclaim {}? [y/N]",
app.marked_count(),
bytesize::ByteSize::b(app.reclaimable_bytes())
);
frame.render_widget(
Paragraph::new(text).block(Block::default().borders(Borders::ALL).title("Confirm")),
area,
);
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use crate::tui::app::App;
use ratatui::backend::TestBackend;
use ratatui::Terminal;
use std::path::PathBuf;
use std::time::{Duration, UNIX_EPOCH};
fn sample_app() -> App {
let files = vec![
FileInfo {
path: PathBuf::from("/tmp/report.pdf").into_boxed_path(),
size: 4_000_000,
modified: UNIX_EPOCH + Duration::from_secs(1_700_000_000),
},
FileInfo {
path: PathBuf::from("/tmp/report-copy.pdf").into_boxed_path(),
size: 4_000_000,
modified: UNIX_EPOCH + Duration::from_secs(1_700_100_000),
},
];
App::from_groups(vec![DuplicateGroup { hash: 0, files }])
}
#[test]
fn draw_renders_without_panic_and_shows_content() {
let app = sample_app();
let params = crate::params::Params::default();
let backend = TestBackend::new(100, 24);
let mut terminal = Terminal::new(backend).unwrap();
terminal.draw(|f| draw(f, &app, &params)).unwrap();
let text: String = terminal
.backend()
.buffer()
.content()
.iter()
.map(|cell| cell.symbol())
.collect();
assert!(text.contains("Groups"));
assert!(text.contains("report"));
assert!(text.contains("marked"));
}
}