13 Commits

Author SHA1 Message Date
Sreedev Kodichath
2a5324b48c [ADD] tui: interactive terminal UI for resolving duplicates
The tool could only act on duplicates non-interactively or through a
line-oriented per-group prompt. Add a full-screen terminal UI (--tui) for
browsing and resolving them visually.

The UI is a two-pane master/detail view: a list of duplicate groups and,
for the selected group, its files with per-file deletion marks. Files can
be marked individually or in bulk by a keep-strategy (newest/oldest/first/
last/shortest/shallowest, reused from the resolver) applied to the current
group or to all groups at once. A footer tracks how many files are marked
and how much space would be reclaimed. Deletion is gated behind an explicit
confirmation, a group always keeps at least one file, and the highlighted
file can be opened in the system default application.

The design keeps the interaction logic pure and independently testable: an
App model translates an abstract key event into an outcome, so navigation,
marking, strategy application, and post-deletion model updates are all unit
tested without a terminal. Rendering (ratatui) and the destructive file I/O
live in the event loop, which restores the terminal on both normal exit and
panic so a crash never leaves it broken. --tui is mutually exclusive with
--keep and --interactive.
2026-07-26 22:02:17 +02:00
Sreedev Kodichath
1442cd0aaf [ADD] cache: on-disk hash cache for faster re-runs
Content hashing is the pipeline's expensive stage; every run re-read and
re-hashed all size-collision candidates from scratch. Persist those hashes
so an unchanged file is skipped on subsequent runs.

Add a path-keyed cache mapping absolute_path -> (size, mtime, strict, hash,
last_seen), stored in a compact hand-rolled little-endian binary format with
a magic+version header. A candidate reuses its stored hash only when size,
mtime, and hash-mode all match; otherwise it is re-hashed and the entry is
refreshed. On save, entries unseen for 30 days are pruned so the file cannot
grow without bound. The hashing stage splits into hash_candidates (parallel,
cache-consulting) and group_hashed (grouping) so lookups stay inside the
parallel pass while cache mutation stays single-threaded after it.

The cache is strictly a performance hint: a missing, corrupt, version-
mismatched, or unwritable cache degrades to a correct full-hash run and
never fails or changes the result. Caching is on by default at the platform
cache directory; --no-cache disables it and --cache-file overrides the path.

The gxhash seed becomes a fixed constant so stored hashes are reproducible
across runs, which is what makes cache hits possible; rand is consequently
no longer a production dependency (now dev-only), and dirs is added for the
platform cache directory.
2026-07-26 20:53:41 +02:00
Sreedev Kodichath
cc7144afbc [ADD] resolver: bulk duplicate resolution via --keep/--force
The tool could only act on duplicates through the default listing or the
manual per-group interactive prompt, with no way to resolve groups in bulk
by a rule. This blocked any scripted or automated use, and bulk resolution
is the top item in the README's proposed operations.

Add a resolver module that reduces each duplicate group to a single kept
file. --keep <newest|oldest|first|last|shortest|shallowest> selects the
keeper: newest/oldest by mtime, first/last and shortest/shallowest by
path, each resolved to a unique keeper through total-order path tiebreaks
so the outcome is deterministic and order-independent. Selection is a pure,
independently tested function; deletion is a separate step.

Deletion is guarded: --keep alone previews (KEEP/DELETE lines and a
would-free summary) and removes nothing, so a mistaken invocation is
harmless. --force performs the deletion, continues past individual
failures, reports freed space, and exits non-zero if any file could not be
removed. clap enforces that --keep excludes --interactive and that --force
requires --keep, so misuse fails before any file is touched.

Also fixes the interactive table's column width, which sized on path
component count instead of path string length so the columns never aligned.
2026-07-26 19:27:49 +02:00
Sreedev Kodichath
3ca931fa8e [REF] core: replace streaming pipeline with staged batch pipeline
The deduplication core was a three-stage streaming producer/consumer
(scan -> group-by-size -> group-by-hash) orchestrated by a Server struct
that ran the stages concurrently on a threadpool. Coordination relied on
AtomicBool flags polled in busy-wait loops, an Arc<Mutex<Vec>> hand-off
queue, DashMap stores, and a per-file Arc<Mutex<FileState>>. Every stage
had to receive its own hand-cloned Arcs with ad-hoc names, and the
busy-wait branches burned a core spinning while waiting for the producer.

Replace it with a staged batch pipeline over owned collections:
pipeline::run(&Params) drives scan -> group_by_size -> group_by_hash in
sequence, with rayon supplying parallelism. With no shared mutable state
between stages, all the Arc plumbing, the AtomicBool coordination, the
Mutex queue, the per-file lock, server.rs, and the threadpool/dashmap
dependencies are gone. FileInfo becomes plain data and each stage is a
pure, independently testable function.

Behavior is preserved except for three authorized deviations: progress
spinners render sequentially rather than concurrently under -p; the
interactive empty-result message prints once instead of twice; and the
interactive "Duplicate Set X of Y" total now counts the confirmed
duplicate groups shown instead of the internal candidate-hash store size.

This lands as the foundation for the cache, CLI, and TUI work that follows.
2026-07-26 14:27:02 +02:00
sreedevk
d102d92bfa cleanup clone littering in server.rs 2025-08-07 23:18:02 +00:00
sreedevk
68c2491fd0 added file store partitioning idea to readme. 2025-07-30 13:08:54 +00:00
sreedevk
da24efbb36 count graphemes instead of chars in path length check 2025-07-30 10:47:58 +00:00
sreedevk
2191763b4d added potential optimizations to readme 2025-07-28 14:22:13 +00:00
sreedevk
b84510f18e tree rendering ui improvements 2025-07-19 16:38:17 +00:00
Sreedev Kodichath
bfb89edd16 Update README.md 2025-07-19 12:32:09 -04:00
sreedevk
21f0eb3cb0 all file type related filtering issues fixed 2025-07-19 16:26:12 +00:00
sreedevk
f5d2d4e22c fix: single file hash groups printed to screen 2025-07-19 14:22:43 +00:00
sreedevk
360af25318 updated readme 2025-07-19 13:57:26 +00:00
19 changed files with 3292 additions and 690 deletions

2
.cargo/config.toml Normal file
View File

@@ -0,0 +1,2 @@
[build]
rustflags = ["-C", "target-feature=+aes,+sse2"]

3
.gitignore vendored
View File

@@ -2,3 +2,6 @@
/Cargo.lock
/result-bin
/.bacon-locations
/docs
/.claude
/.superpowers

1383
Cargo.lock generated

File diff suppressed because it is too large Load Diff

View File

@@ -1,6 +1,6 @@
[package]
name = "deduplicator"
version = "0.3.1"
version = "0.3.2"
edition = "2021"
description = "find,filter and delete duplicate files"
repository = "https://github.com/sreedevk/deduplicator"
@@ -21,16 +21,17 @@ anyhow = "1.0.68"
bytesize = "2.0.1"
chrono = "0.4.23"
clap = { version = "4.0.32", features = ["derive"] }
dashmap = { version = "6.1.0", features = ["rayon"] }
dirs = "6.0.0"
globwalk = "0.9.1"
gxhash = { version = "3.4.1", default-features = false }
indicatif = { version = "0.18.0", features = ["rayon"] }
memmap2 = "0.9.7"
open = "5.4.0"
pathdiff = "0.2.1"
prettytable-rs = "0.10.0"
rand = "0.9.1"
ratatui = "0.30.2"
rayon = "1.6.1"
threadpool = "1.8.1"
unicode-segmentation = "1.12.0"
[profile.release]
strip = true
@@ -56,3 +57,4 @@ cargo-dist-version = "0.0.7"
[dev-dependencies]
tempfile = "3.20.0"
rand = "0.9.1"

View File

@@ -50,7 +50,8 @@ deduplicator ~/Media --min-size 100mb
```
## Demo
![record](https://github.com/user-attachments/assets/fcfdd9bf-4d05-41b6-a82e-367634eeaa73)
![demo](https://github.com/user-attachments/assets/bdb95831-542d-4902-a458-4e0f5d171a33)
@@ -80,6 +81,12 @@ $ RUSTFLAGS="-C target-cpu=native" cargo install deduplicator --git https://gith
$ RUSTFLAGS="-C target-feature=+aes,+sse2" cargo install --git https://github.com/sreedevk/deduplicator
```
### Manual Installation
- Download the right pre-compiled binary archive for your platform from [github release page](https://github.com/sreedevk/deduplicator/releases/tag/latest).
- Decompress it using `tar -zxvf <archive>.tar.gz`
- Move it to a directory included in `$PATH`.
- ideally `/usr/local/bin/`.
## Performance
Deduplicator uses size comparison and [GxHash](https://docs.rs/gxhash/latest/gxhash/) to quickly check a large number of files to find duplicates. its also heavily parallelized. The default behavior of deduplicator is to only hash the first page (4K) of the file. This is to ensure that performance is the default priority. You can modify this behavior by using the `--strict` flag which will hash the whole file and ensure that 2 files are indeed duplicates. I'll add benchmarks in future versions.
@@ -149,6 +156,8 @@ dust 'bench_artifacts'
## proposed
- [ ] parallelization
- [ ] scanning + processing sw + processing hw + formatting + printing
- [ ] user supplied cache file path for faster re-runs
- [ ] hardlinks / symlinks support
- [ ] max file path size should use the last set of duplicates
- [ ] add more unit tests
- [ ] test against different filesystems
@@ -160,6 +169,7 @@ dust 'bench_artifacts'
- [ ] tui
- [ ] change the default hashing method to include the first & last page of a file (8K)
- [ ] provide option to localize duplicate detection to arbitrary levels relative to current directory
- [ ] localize file meta store locks to sub path levels to avoid global lock contention from multiple threads.
- [ ] bulk operations
- [ ] --keep-latest
- [ ] --keep-oldest
@@ -169,6 +179,13 @@ dust 'bench_artifacts'
- [ ] fix: partial hash collision - a file full of null bytes ("\0") and an empty file. This is a known trade off in gxhash.
- [ ] include initial pages and final pages of the file
- [ ] append the offset between the last initial page hashed and the first final page hashed in the content passed to the hasher.
- [ ] potential optimizations
- [ ] lookup the memory efficiency gains if instead of directly inserting into a hashmap, deduplicator looksup the file in a bloom filter. this way, the duplicate store hashmap does
not require to be locked for every single file. The bloom filter can be stored in an atomically updateable type to improve performance as well.
## v0.3.2
- [x] fix: single file groups are printed to screen
- [x] fix: --exclude-types and --types flag behave identically.
## v0.3.1
- [x] parallelization

249
src/cache.rs Normal file
View File

@@ -0,0 +1,249 @@
use std::collections::HashMap;
use std::fs;
use std::path::{Path, PathBuf};
use std::time::{SystemTime, UNIX_EPOCH};
const MAGIC: &[u8; 4] = b"DDUP";
const VERSION: u8 = 1;
const RECORD_FIXED_LEN: usize = 45;
pub const CACHE_TTL_DAYS: u64 = 30;
pub struct CacheEntry {
pub size: u64,
pub mtime: u64,
pub strict: bool,
pub hash: u128,
pub last_seen: u64,
}
pub struct Cache {
entries: HashMap<PathBuf, CacheEntry>,
enabled: bool,
}
pub fn mtime_nanos(modified: SystemTime) -> u64 {
modified
.duration_since(UNIX_EPOCH)
.map(|d| d.as_nanos() as u64)
.unwrap_or(0)
}
impl Cache {
pub fn disabled() -> Self {
Self {
entries: HashMap::new(),
enabled: false,
}
}
fn empty_enabled() -> Self {
Self {
entries: HashMap::new(),
enabled: true,
}
}
pub fn load(path: &Path) -> Self {
match fs::read(path) {
Ok(bytes) => Self::parse(&bytes).unwrap_or_else(Self::empty_enabled),
Err(_) => Self::empty_enabled(),
}
}
fn parse(bytes: &[u8]) -> Option<Self> {
if bytes.len() < 5 || &bytes[0..4] != MAGIC || bytes[4] != VERSION {
return None;
}
let mut entries = HashMap::new();
let mut cur = &bytes[5..];
while cur.len() >= RECORD_FIXED_LEN {
let hash = u128::from_le_bytes(cur[0..16].try_into().ok()?);
let size = u64::from_le_bytes(cur[16..24].try_into().ok()?);
let mtime = u64::from_le_bytes(cur[24..32].try_into().ok()?);
let last_seen = u64::from_le_bytes(cur[32..40].try_into().ok()?);
let strict = cur[40] != 0;
let path_len = u32::from_le_bytes(cur[41..45].try_into().ok()?) as usize;
let rest = &cur[RECORD_FIXED_LEN..];
if rest.len() < path_len {
break;
}
match std::str::from_utf8(&rest[..path_len]) {
Ok(s) => {
entries.insert(
PathBuf::from(s),
CacheEntry {
size,
mtime,
strict,
hash,
last_seen,
},
);
}
Err(_) => {}
}
cur = &rest[path_len..];
}
Some(Self {
entries,
enabled: true,
})
}
pub fn lookup(&self, path: &Path, size: u64, mtime: u64, strict: bool) -> Option<u128> {
let entry = self.entries.get(path)?;
match entry.size == size && entry.mtime == mtime && entry.strict == strict {
true => Some(entry.hash),
false => None,
}
}
pub fn record(
&mut self,
path: &Path,
size: u64,
mtime: u64,
strict: bool,
hash: u128,
now_secs: u64,
) {
if path.to_str().is_none() {
return;
}
self.entries.insert(
path.to_path_buf(),
CacheEntry {
size,
mtime,
strict,
hash,
last_seen: now_secs,
},
);
}
pub fn save(&self, path: &Path, ttl_secs: u64, now_secs: u64) {
if !self.enabled {
return;
}
let mut buf: Vec<u8> = Vec::new();
buf.extend_from_slice(MAGIC);
buf.push(VERSION);
for (entry_path, entry) in &self.entries {
if now_secs.saturating_sub(entry.last_seen) > ttl_secs {
continue;
}
let path_str = match entry_path.to_str() {
Some(s) => s,
None => continue,
};
buf.extend_from_slice(&entry.hash.to_le_bytes());
buf.extend_from_slice(&entry.size.to_le_bytes());
buf.extend_from_slice(&entry.mtime.to_le_bytes());
buf.extend_from_slice(&entry.last_seen.to_le_bytes());
buf.push(entry.strict as u8);
buf.extend_from_slice(&(path_str.len() as u32).to_le_bytes());
buf.extend_from_slice(path_str.as_bytes());
}
if let Some(parent) = path.parent() {
let _ = fs::create_dir_all(parent);
}
if let Err(err) = fs::write(path, &buf) {
eprintln!("warning: failed to write cache to {}: {err}", path.display());
}
}
pub fn default_path() -> Option<PathBuf> {
dirs::cache_dir().map(|dir| dir.join("deduplicator").join("cache.bin"))
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::path::Path;
use std::time::{Duration, UNIX_EPOCH};
use tempfile::TempDir;
#[test]
fn lookup_hits_on_exact_match() {
let mut c = Cache::disabled();
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), Some(42));
}
#[test]
fn lookup_misses_on_any_field_change_or_absent() {
let mut c = Cache::disabled();
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
assert_eq!(c.lookup(Path::new("/a/b"), 101, 200, false), None);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 201, false), None);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, true), None);
assert_eq!(c.lookup(Path::new("/a/x"), 100, 200, false), None);
}
#[test]
fn save_then_load_roundtrips_entries() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
let mut c = Cache::load(&file);
c.record(Path::new("/a/b"), 100, 200, false, 42, 1000);
c.record(Path::new("/c/d"), 5, 6, true, 99, 1000);
c.save(&file, 86_400, 1000);
let loaded = Cache::load(&file);
assert_eq!(loaded.lookup(Path::new("/a/b"), 100, 200, false), Some(42));
assert_eq!(loaded.lookup(Path::new("/c/d"), 5, 6, true), Some(99));
}
#[test]
fn load_returns_empty_on_corrupt_file() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
std::fs::write(&file, b"not a valid cache file at all").unwrap();
let c = Cache::load(&file);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), None);
}
#[test]
fn load_returns_empty_on_version_mismatch() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
std::fs::write(&file, b"DDUP\x02").unwrap();
let c = Cache::load(&file);
assert_eq!(c.lookup(Path::new("/a/b"), 100, 200, false), None);
}
#[test]
fn save_prunes_entries_older_than_ttl() {
let dir = TempDir::new().unwrap();
let file = dir.path().join("cache.bin");
let mut c = Cache::load(&file);
c.record(Path::new("/old"), 1, 2, false, 10, 1000);
c.record(Path::new("/new"), 3, 4, false, 20, 5000);
c.save(&file, 1000, 5000);
let loaded = Cache::load(&file);
assert_eq!(loaded.lookup(Path::new("/old"), 1, 2, false), None);
assert_eq!(loaded.lookup(Path::new("/new"), 3, 4, false), Some(20));
}
#[test]
fn mtime_nanos_converts_systemtime() {
let t = UNIX_EPOCH + Duration::from_nanos(1234);
assert_eq!(mtime_nanos(t), 1234);
}
}

View File

@@ -5,22 +5,14 @@ use std::{
fs,
io::Read,
path::{Path, PathBuf},
sync::{Arc, Mutex},
time::SystemTime,
};
#[derive(Debug, Clone, PartialEq)]
pub enum FileState {
Unprocessed,
SwProcessed,
}
#[derive(Debug, Clone)]
pub struct FileInfo {
pub path: Box<Path>,
pub size: u64,
pub modified: SystemTime,
pub state: Arc<Mutex<FileState>>,
}
impl FileInfo {
@@ -53,19 +45,8 @@ impl FileInfo {
path: path.into_boxed_path(),
size: filemeta.len(),
modified: filemeta.modified()?,
state: Arc::new(Mutex::new(FileState::Unprocessed)),
})
}
pub fn sw_processed(&self) {
let mut self_state = self.state.lock().unwrap();
*self_state = FileState::SwProcessed;
}
pub fn is_sw_processed(&self) -> bool {
let self_state = self.state.lock().unwrap();
*self_state == FileState::SwProcessed
}
}
#[cfg(test)]

View File

@@ -1,24 +1,23 @@
use crate::{fileinfo::FileInfo, params::Params};
use crate::{fileinfo::FileInfo, params::Params, pipeline::DedupReport};
use anyhow::Result;
use chrono::{DateTime, Utc};
use dashmap::DashMap;
use pathdiff::diff_paths;
use rayon::prelude::*;
use std::{path::PathBuf, sync::Arc};
use std::path::PathBuf;
const YELLOW: &str = "\x1b[33m";
const RESET: &str = "\x1b[0m";
pub struct Formatter;
impl Formatter {
pub fn human_path(file: &FileInfo, aargs: &Params, min_path_length: usize) -> Result<String> {
pub fn human_path(file: &FileInfo, aargs: &Params, max_path_length: usize) -> Result<String> {
let base_directory: PathBuf = aargs.get_directory()?;
let relative_path = diff_paths(&file.path, base_directory).unwrap_or_default();
let formatted_path = format!(
"{:<0width$}",
relative_path.to_str().unwrap_or_default().to_string(),
width = min_path_length
width = max_path_length
);
Ok(formatted_path)
@@ -33,32 +32,39 @@ impl Formatter {
Ok(modified_time.format("%Y-%m-%d %H:%M:%S").to_string())
}
pub fn print(raw: Arc<DashMap<u128, Vec<FileInfo>>>, max_path_len: u64, aargs: &Params) {
print!("{}", "\n".repeat(if aargs.progress { 2 } else { 1 })); // spacing
pub fn print(report: &DedupReport, aargs: &Params) {
print!("{}", "\n".repeat(if aargs.progress { 2 } else { 1 }));
if raw.is_empty() {
if report.groups.is_empty() {
println!("No duplicates found matching your search criteria.");
} else {
raw.par_iter().for_each(|sref| {
let mut ostring = format!("{}{:32x}{}\n", YELLOW, sref.key(), RESET);
let subfields = sref
.value()
.par_iter()
.map(|finfo| {
format!(
"├─ {}\t{}\t{}\n",
Self::human_path(finfo, aargs, max_path_len as usize)
.expect("path formatting failed."),
Self::human_filesize(finfo).expect("filesize formatting failed."),
Self::human_mtime(finfo).expect("modified time formatting failed.")
)
})
.collect::<String>();
ostring.push_str(&subfields);
println!("{ostring}");
});
return;
}
report.groups.par_iter().for_each(|group| {
let mut ostring = format!("{}{:32x}{}\n", YELLOW, group.hash, RESET);
let subfields = group
.files
.par_iter()
.enumerate()
.map(|(i, finfo)| {
let nodechar = if i == group.files.len() - 1 {
"└─"
} else {
"├─"
};
format!(
"{}\t{}\t{}\t{}\n",
nodechar,
Self::human_path(finfo, aargs, report.max_path_len)
.expect("path formatting failed."),
Self::human_filesize(finfo).expect("filesize formatting failed."),
Self::human_mtime(finfo).expect("modified time formatting failed.")
)
})
.collect::<String>();
ostring.push_str(&subfields);
println!("{ostring}");
});
}
}

View File

@@ -1,45 +1,42 @@
use crate::{fileinfo::FileInfo, formatter::Formatter, params::Params};
use crate::{fileinfo::FileInfo, formatter::Formatter, params::Params, pipeline::DuplicateGroup};
use anyhow::Result;
use dashmap::DashMap;
use prettytable::{format, row, Table};
use std::{
io::{self, Write},
sync::Arc,
};
use std::io::{self, Write};
use unicode_segmentation::UnicodeSegmentation;
pub struct Interactive;
impl Interactive {
pub fn init(result: Arc<DashMap<u128, Vec<FileInfo>>>, app_args: &Params) -> Result<()> {
result
.clone()
.iter()
.filter(|i| i.value().len() > 1)
.enumerate()
.for_each(|(gindex, i)| {
let group = i.value();
let mut itable = Table::new();
itable.set_format(*format::consts::FORMAT_NO_BORDER_LINE_SEPARATOR);
itable.set_titles(row!["index", "filename", "size", "updated_at"]);
pub fn init(groups: &[DuplicateGroup], app_args: &Params) -> Result<()> {
if groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
}
let max_path_size = group
.iter()
.map(|f| f.path.iter().count())
.max()
.unwrap_or_default();
groups.iter().enumerate().for_each(|(gindex, group)| {
let files = &group.files;
let mut itable = Table::new();
itable.set_format(*format::consts::FORMAT_NO_BORDER_LINE_SEPARATOR);
itable.set_titles(row!["index", "filename", "size", "updated_at"]);
group.iter().enumerate().for_each(|(index, file)| {
itable.add_row(row![
index,
Formatter::human_path(file, app_args, max_path_size).unwrap_or_default(),
Formatter::human_filesize(file).unwrap_or_default(),
Formatter::human_mtime(file).unwrap_or_default()
]);
});
let max_path_size = files
.iter()
.map(|f| f.path.to_string_lossy().graphemes(true).count())
.max()
.unwrap_or_default();
Self::process_group_action(group, gindex, result.len(), itable);
files.iter().enumerate().for_each(|(index, file)| {
itable.add_row(row![
index,
Formatter::human_path(file, app_args, max_path_size).unwrap_or_default(),
Formatter::human_filesize(file).unwrap_or_default(),
Formatter::human_mtime(file).unwrap_or_default()
]);
});
Self::process_group_action(files, gindex, groups.len(), itable);
});
Ok(())
}

View File

@@ -1,35 +1,29 @@
mod cache;
mod fileinfo;
mod formatter;
mod interactive;
mod params;
mod pipeline;
mod processor;
mod resolver;
mod scanner;
mod server;
mod tui;
use self::{formatter::Formatter, interactive::Interactive, server::Server};
use self::{formatter::Formatter, interactive::Interactive};
use anyhow::Result;
use clap::Parser;
use params::Params;
use std::sync::atomic::Ordering;
fn main() -> Result<()> {
let app_args = Params::parse();
let server = Server::new(app_args.clone());
let params = Params::parse();
let report = pipeline::run(&params)?;
server.start()?;
match app_args.interactive {
false => {
Formatter::print(
server.hw_duplicate_set,
server.max_file_path_len.load(Ordering::Acquire),
&app_args,
);
}
true => {
Interactive::init(server.hw_duplicate_set, &app_args)?;
}
};
match (params.tui, params.keep, params.interactive) {
(true, _, _) => tui::run(report, &params)?,
(false, Some(strategy), _) => resolver::run(&report, strategy, params.force, &params)?,
(false, None, true) => Interactive::init(&report.groups, &params)?,
(false, None, false) => Formatter::print(&report, &params),
}
Ok(())
}

View File

@@ -2,7 +2,6 @@ use std::{fs, path::PathBuf};
use anyhow::Result;
use clap::{Parser, ValueHint};
use std::collections::HashSet;
#[derive(Parser, Debug, Default, Clone)]
#[command(author, version, about, long_about = None)]
@@ -19,6 +18,12 @@ pub struct Params {
/// Delete files interactively
#[arg(long, short)]
pub interactive: bool,
/// Keep one file per duplicate group by this rule and remove the rest
#[arg(long, conflicts_with = "interactive")]
pub keep: Option<crate::resolver::KeepStrategy>,
/// Actually delete the duplicates (without this, --keep only previews)
#[arg(long, visible_alias = "yes", requires = "keep")]
pub force: bool,
/// Minimum filesize of duplicates to scan (e.g., 100B/1K/2M/3G/4T).
#[arg(long, short = 'm', default_value = "1b")]
pub min_size: Option<String>,
@@ -37,6 +42,15 @@ pub struct Params {
/// Show Progress spinners & metrics
#[arg(long, short = 'p', default_value = "false")]
pub progress: bool,
/// Disable the on-disk hash cache
#[arg(long)]
pub no_cache: bool,
/// Use a specific cache file instead of the default location
#[arg(long, value_name = "PATH")]
pub cache_file: Option<PathBuf>,
/// Browse and resolve duplicates in an interactive terminal UI
#[arg(long, conflicts_with_all = ["keep", "interactive"])]
pub tui: bool,
}
impl Params {
@@ -56,48 +70,4 @@ impl Params {
let dir = fs::canonicalize(dir_path)?;
Ok(dir)
}
pub fn types_intersection(itypes: &str, xtypes: &str) -> String {
let iset = itypes
.split(",")
.map(String::from)
.collect::<HashSet<String>>();
let xset = xtypes
.split(",")
.map(String::from)
.collect::<HashSet<String>>();
iset.difference(&xset)
.cloned()
.collect::<Vec<String>>()
.join(",")
}
pub fn get_types(&self) -> Option<String> {
match &self.types {
Some(itypes) => match &self.exclude_types {
Some(xtypes) => Some(Self::types_intersection(itypes, xtypes)),
None => Some(itypes.to_string()),
},
None => self.exclude_types.as_ref().map(|xtypes| xtypes.to_string()),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn mixing_include_and_exclude_types_works_as_expected() {
let params = Params {
types: Some(String::from("js,xml,ts,pdf,tiff")),
exclude_types: Some(String::from("js,ts,xml")),
..Default::default()
};
assert!(params
.get_types()
.is_some_and(|x| x == "pdf,tiff" || x == "tiff,pdf"))
}
}

182
src/pipeline.rs Normal file
View File

@@ -0,0 +1,182 @@
use anyhow::Result;
use indicatif::{ProgressBar, ProgressStyle};
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use unicode_segmentation::UnicodeSegmentation;
use crate::cache::{mtime_nanos, Cache, CACHE_TTL_DAYS};
use crate::fileinfo::FileInfo;
use crate::params::Params;
use crate::processor;
use crate::scanner::Scanner;
const HASH_SEED: i64 = 0x00DE_D0CA_C4E5_EED1;
fn now_secs() -> u64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs())
.unwrap_or(0)
}
pub struct DuplicateGroup {
pub hash: u128,
pub files: Vec<FileInfo>,
}
pub struct DedupReport {
pub groups: Vec<DuplicateGroup>,
pub max_path_len: usize,
}
pub(crate) fn spinner(enabled: bool, message: &'static str) -> ProgressBar {
let bar = if enabled {
ProgressBar::new_spinner()
} else {
ProgressBar::hidden()
};
let style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")
.expect("valid progress template");
bar.set_style(style);
bar.enable_steady_tick(Duration::from_millis(50));
bar.set_message(message);
bar
}
pub fn run(params: &Params) -> Result<DedupReport> {
let seed = HASH_SEED;
let cache_path = match params.no_cache {
true => None,
false => params.cache_file.clone().or_else(Cache::default_path),
};
let mut cache = match (params.no_cache, &cache_path) {
(false, Some(path)) => Cache::load(path),
_ => Cache::disabled(),
};
let files = Scanner::new(params)?.scan()?;
let size_groups = processor::group_by_size(files, params.progress);
let candidates: Vec<FileInfo> = size_groups
.into_iter()
.filter(|group| group.len() > 1)
.flatten()
.collect();
let max_path_len = candidates
.iter()
.map(|file| file.path.to_string_lossy().graphemes(true).count())
.max()
.unwrap_or(0);
let hashed = processor::hash_candidates(candidates, params.strict, seed, params.progress, &cache);
let now = now_secs();
for (hash, file) in &hashed {
let mtime = mtime_nanos(file.modified);
cache.record(&file.path, file.size, mtime, params.strict, *hash, now);
}
let groups = processor::group_hashed(hashed);
if let Some(path) = &cache_path {
cache.save(path, CACHE_TTL_DAYS * 86_400, now);
}
Ok(DedupReport {
groups,
max_path_len,
})
}
#[cfg(test)]
mod tests {
use super::run;
use crate::params::Params;
use anyhow::Result;
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
#[test]
fn run_finds_a_single_group_of_identical_files() -> Result<()> {
let root = TempDir::new()?;
let duplicate = vec![7u8; 200_000];
for name in ["a.bin", "b.bin"] {
let mut file = File::create_new(root.path().join(name))?;
file.write_all(&duplicate)?;
}
let mut unique = File::create_new(root.path().join("c.bin"))?;
unique.write_all(&vec![9u8; 100_000])?;
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
let report = run(&params)?;
assert_eq!(report.groups.len(), 1);
assert_eq!(report.groups[0].files.len(), 2);
Ok(())
}
}
#[cfg(test)]
mod cache_tests {
use super::run;
use crate::params::Params;
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
fn make_tree(root: &TempDir) {
for name in ["a.bin", "b.bin"] {
let mut f = File::create_new(root.path().join(name)).unwrap();
f.write_all(&vec![7u8; 200_000]).unwrap();
}
let mut u = File::create_new(root.path().join("c.bin")).unwrap();
u.write_all(&vec![9u8; 100_000]).unwrap();
}
#[test]
fn run_twice_with_cache_is_identical_and_writes_the_file() {
let root = TempDir::new().unwrap();
make_tree(&root);
let cache_file = root.path().join("dd.cache");
let params = Params {
dir: Some(root.path().into()),
cache_file: Some(cache_file.clone()),
..Default::default()
};
let first = run(&params).unwrap();
assert!(cache_file.exists());
let second = run(&params).unwrap();
assert_eq!(first.groups.len(), second.groups.len());
assert_eq!(first.groups.len(), 1);
assert_eq!(second.groups[0].files.len(), 2);
}
#[test]
fn no_cache_does_not_write_a_cache_file() {
let root = TempDir::new().unwrap();
make_tree(&root);
let cache_file = root.path().join("dd.cache");
let params = Params {
dir: Some(root.path().into()),
no_cache: true,
cache_file: Some(cache_file.clone()),
..Default::default()
};
run(&params).unwrap();
assert!(!cache_file.exists());
}
}

View File

@@ -1,381 +1,214 @@
use anyhow::Result;
use dashmap::DashMap;
use indicatif::{MultiProgress, ProgressBar, ProgressStyle};
use rayon::iter::IntoParallelRefMutIterator;
use std::collections::HashMap;
use rayon::prelude::{IntoParallelIterator, ParallelIterator};
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::{Arc, Mutex, TryLockError, TryLockResult};
use std::time::Duration;
use crate::cache::{mtime_nanos, Cache};
use crate::fileinfo::FileInfo;
use crate::params::Params;
use crate::pipeline::{spinner, DuplicateGroup};
pub struct Processor {}
pub fn group_by_size(files: Vec<FileInfo>, progress: bool) -> Vec<Vec<FileInfo>> {
let bar = spinner(progress, "files grouped by size");
impl Processor {
pub fn hashwise(
app_args: Arc<Params>,
sw_store: Arc<DashMap<u64, Vec<FileInfo>>>,
hw_store: Arc<DashMap<u128, Vec<FileInfo>>>,
progress_bar_box: Arc<MultiProgress>,
max_file_size: Arc<AtomicU64>,
seed: i64,
sw_sorting_finished: Arc<AtomicBool>,
) -> Result<()> {
let progress_bar = match app_args.progress {
true => progress_bar_box.add(ProgressBar::new_spinner()),
false => ProgressBar::hidden(),
};
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")?;
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("files grouped by hash.");
loop {
let keys: Vec<u64> = sw_store
.clone()
.iter()
.filter(|i| !i.value().iter().all(|x| x.is_sw_processed()))
.filter(|i| i.value().len() > 1)
.map(|i| *i.key())
.collect();
if keys.is_empty() {
match sw_sorting_finished.load(std::sync::atomic::Ordering::Relaxed) {
true => {
progress_bar.finish_with_message("files grouped by hash.");
break Ok(());
}
false => continue,
}
} else {
keys.into_par_iter().for_each(|key| {
let mut group: Vec<FileInfo> = sw_store.get(&key).unwrap().to_vec();
if group.len() > 1 {
group.par_iter_mut().for_each(|file| {
progress_bar.inc(1);
file.sw_processed();
let fhash = match app_args.strict {
true => file.hash(seed).expect("hashing file failed."),
false => file.initpages_hash(seed).expect("hashing file failed."),
};
Self::compare_and_update_max_path_len(
max_file_size.clone(),
file.path.to_string_lossy().len() as u64,
);
hw_store
.entry(fhash)
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file.clone()]);
});
};
});
}
}
let mut buckets: HashMap<u64, Vec<FileInfo>> = HashMap::new();
for file in files {
bar.inc(1);
buckets.entry(file.size).or_default().push(file);
}
pub fn compare_and_update_max_path_len(current: Arc<AtomicU64>, next: u64) {
if current.load(Ordering::Relaxed) < next {
current.store(next, Ordering::Release);
}
}
bar.finish_with_message("files grouped by size");
buckets.into_values().collect()
}
pub fn sizewise(
app_args: Arc<Params>,
scanner_finished: Arc<AtomicBool>,
store: Arc<DashMap<u64, Vec<FileInfo>>>,
files: Arc<Mutex<Vec<FileInfo>>>,
progress_bar_box: Arc<MultiProgress>,
) -> Result<()> {
let progress_bar = match app_args.progress {
true => progress_bar_box.add(ProgressBar::new_spinner()),
false => ProgressBar::hidden(),
};
pub fn hash_candidates(
candidates: Vec<FileInfo>,
strict: bool,
seed: i64,
progress: bool,
cache: &Cache,
) -> Vec<(u128, FileInfo)> {
let bar = spinner(progress, "files grouped by hash");
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")?;
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("files grouped by size");
loop {
let fileopt: Option<FileInfo> = {
match files.try_lock() {
Ok(mut flist) => flist.pop(),
TryLockResult::Err(TryLockError::WouldBlock) => None,
_ => None,
}
};
match fileopt {
Some(file) => {
progress_bar.inc(1);
store
.entry(file.size)
.and_modify(|fileset| fileset.push(file.clone()))
.or_insert_with(|| vec![file]);
continue;
}
None => match scanner_finished.load(std::sync::atomic::Ordering::Relaxed) {
true => {
progress_bar.finish_with_message("files grouped by size");
break Ok(());
}
false => continue,
let hashed: Vec<(u128, FileInfo)> = candidates
.into_par_iter()
.map(|file| {
bar.inc(1);
let mtime = mtime_nanos(file.modified);
let hash = match cache.lookup(&file.path, file.size, mtime, strict) {
Some(cached) => cached,
None => match strict {
true => file.hash(seed).expect("hashing file failed."),
false => file.initpages_hash(seed).expect("hashing file failed."),
},
}
}
};
(hash, file)
})
.collect();
bar.finish_with_message("files grouped by hash.");
hashed
}
pub fn group_hashed(hashed: Vec<(u128, FileInfo)>) -> Vec<DuplicateGroup> {
let mut buckets: HashMap<u128, Vec<FileInfo>> = HashMap::new();
for (hash, file) in hashed {
buckets.entry(hash).or_default().push(file);
}
buckets
.into_iter()
.filter(|(_, files)| files.len() > 1)
.map(|(hash, files)| DuplicateGroup { hash, files })
.collect()
}
#[cfg(test)]
mod tests {
mod staged_tests {
use anyhow::Result;
use dashmap::DashMap;
use indicatif::MultiProgress;
use rand::Rng;
use std::fs::File;
use std::io::Write;
use std::sync::atomic::{AtomicBool, AtomicU64};
use std::sync::{Arc, Mutex};
use tempfile::TempDir;
use crate::{fileinfo::FileInfo, params::Params};
use super::Processor;
use crate::fileinfo::FileInfo;
fn generate_bytes(size: usize) -> Vec<u8> {
let mut rng = rand::rng();
(0..size).map(|_| rng.random::<u8>()).collect::<Vec<u8>>()
}
fn write_files(root: &TempDir, specs: Vec<(&str, Vec<u8>)>) -> Result<Vec<FileInfo>> {
specs
.into_iter()
.map(|(name, content)| {
let path = root.path().join(name);
let mut file = File::create_new(&path)?;
file.write_all(&content)?;
FileInfo::new(path)
})
.collect()
}
#[test]
fn hashwise_sorting_two_files_with_identical_init_pages_only_strict_mode() -> Result<()> {
fn group_by_size_separates_files_of_different_sizes() -> Result<()> {
let root = TempDir::new()?;
let content = generate_bytes(16384);
let mut content_x = content.clone();
let mut content_y = content.clone();
content_x.extend(generate_bytes(1720320));
content_y.extend(generate_bytes(1720320));
let files = [
(root.path().join("fileone.bin"), content_x),
(root.path().join("filetwo.bin"), content_y),
];
for (fpath, content) in files.iter() {
let mut f = File::create_new(fpath)?;
f.write_all(content)?;
}
let dupstore = Arc::new(DashMap::new());
let file_queue = Arc::new(Mutex::new(
files
.iter()
.map(|f| FileInfo::new(f.0.clone()).unwrap())
.collect::<Vec<FileInfo>>(),
));
let hw_dupstore = Arc::new(DashMap::new());
Processor::sizewise(
Arc::new(Params::default()),
Arc::new(AtomicBool::new(true)),
dupstore.clone(),
file_queue,
Arc::new(MultiProgress::new()),
let files = write_files(
&root,
vec![
("fileone.bin", generate_bytes(282624)),
("filetwo.bin", generate_bytes(1720320)),
],
)?;
let args = Params {
strict: true,
..Default::default()
};
Processor::hashwise(
Arc::new(args),
dupstore.clone(),
hw_dupstore.clone(),
Arc::new(MultiProgress::new()),
Arc::new(AtomicU64::new(32)),
300,
Arc::new(AtomicBool::new(true)),
)?;
assert_eq!(hw_dupstore.len(), 2);
let groups = super::group_by_size(files, false);
assert_eq!(groups.len(), 2);
Ok(())
}
#[test]
fn hashwise_sorting_two_files_with_identical_init_pages_only_fast_mode() -> Result<()> {
fn group_by_size_buckets_same_size_files_together() -> Result<()> {
let root = TempDir::new()?;
let content = generate_bytes(16384);
let mut content_x = content.clone();
let mut content_y = content.clone();
content_x.extend(generate_bytes(1720320));
content_y.extend(generate_bytes(1720320));
let files = [
(root.path().join("fileone.bin"), content_x),
(root.path().join("filetwo.bin"), content_y),
];
for (fpath, content) in files.iter() {
let mut f = File::create_new(fpath)?;
f.write_all(content)?;
}
let dupstore = Arc::new(DashMap::new());
let file_queue = Arc::new(Mutex::new(
files
.iter()
.map(|f| FileInfo::new(f.0.clone()).unwrap())
.collect::<Vec<FileInfo>>(),
));
let hw_dupstore = Arc::new(DashMap::new());
Processor::sizewise(
Arc::new(Params::default()),
Arc::new(AtomicBool::new(true)),
dupstore.clone(),
file_queue,
Arc::new(MultiProgress::new()),
let files = write_files(
&root,
vec![
("fileone.bin", generate_bytes(282624)),
("filetwo.bin", generate_bytes(282624)),
],
)?;
Processor::hashwise(
Arc::new(Params::default()),
dupstore.clone(),
hw_dupstore.clone(),
Arc::new(MultiProgress::new()),
Arc::new(AtomicU64::new(32)),
300,
Arc::new(AtomicBool::new(true)),
)?;
assert_eq!(hw_dupstore.len(), 1);
let groups = super::group_by_size(files, false);
assert_eq!(groups.len(), 1);
Ok(())
}
#[test]
fn hashwise_sorting_two_files_with_identical_data() -> Result<()> {
fn group_by_hash_fast_mode_matches_identical_init_pages() -> Result<()> {
let root = TempDir::new()?;
let shared = generate_bytes(16384);
let mut content_x = shared.clone();
let mut content_y = shared.clone();
content_x.extend(generate_bytes(1720320));
content_y.extend(generate_bytes(1720320));
let files = write_files(
&root,
vec![("fileone.bin", content_x), ("filetwo.bin", content_y)],
)?;
let groups = super::group_hashed(super::hash_candidates(files, false, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 1);
Ok(())
}
#[test]
fn group_by_hash_strict_mode_rejects_different_tails() -> Result<()> {
let root = TempDir::new()?;
let shared = generate_bytes(16384);
let mut content_x = shared.clone();
let mut content_y = shared.clone();
content_x.extend(generate_bytes(1720320));
content_y.extend(generate_bytes(1720320));
let files = write_files(
&root,
vec![("fileone.bin", content_x), ("filetwo.bin", content_y)],
)?;
let groups = super::group_hashed(super::hash_candidates(files, true, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 0);
Ok(())
}
#[test]
fn group_by_hash_matches_identical_files() -> Result<()> {
let root = TempDir::new()?;
let content = generate_bytes(282624);
let files = [
(root.path().join("fileone.bin"), content.clone()),
(root.path().join("filetwo.bin"), content.clone()),
];
for (fpath, content) in files.iter() {
let mut f = File::create_new(fpath)?;
f.write_all(content)?;
}
let dupstore = Arc::new(DashMap::new());
let file_queue = Arc::new(Mutex::new(
files
.iter()
.map(|f| FileInfo::new(f.0.clone()).unwrap())
.collect::<Vec<FileInfo>>(),
));
let hw_dupstore = Arc::new(DashMap::new());
Processor::sizewise(
Arc::new(Params::default()),
Arc::new(AtomicBool::new(true)),
dupstore.clone(),
file_queue,
Arc::new(MultiProgress::new()),
let files = write_files(
&root,
vec![
("fileone.bin", content.clone()),
("filetwo.bin", content.clone()),
],
)?;
Processor::hashwise(
Arc::new(Params::default()),
dupstore.clone(),
hw_dupstore.clone(),
Arc::new(MultiProgress::new()),
Arc::new(AtomicU64::new(32)),
300,
Arc::new(AtomicBool::new(true)),
)?;
assert_eq!(hw_dupstore.len(), 1);
let groups = super::group_hashed(super::hash_candidates(files, false, 300, false, &crate::cache::Cache::disabled()));
assert_eq!(groups.len(), 1);
Ok(())
}
#[test]
fn sizewise_sorting_two_files_of_different_sizes() -> Result<()> {
let root = TempDir::new()?;
let files = [
(root.path().join("fileone.bin"), generate_bytes(282624)),
(root.path().join("filetwo.bin"), generate_bytes(1720320)),
];
fn hash_candidates_uses_cached_hash_when_valid() {
let root = TempDir::new().unwrap();
let path = root.path().join("f.bin");
let mut f = File::create_new(&path).unwrap();
f.write_all(b"real content for cache hit test").unwrap();
let info = FileInfo::new(path.clone()).unwrap();
let mtime = crate::cache::mtime_nanos(info.modified);
for (fpath, content) in files.iter() {
let mut f = File::create_new(fpath)?;
f.write_all(content)?;
}
let mut cache = crate::cache::Cache::disabled();
cache.record(&path, info.size, mtime, false, 0xDEAD_BEEF, 0);
let file_queue = Arc::new(Mutex::new(
files
.iter()
.map(|f| FileInfo::new(f.0.clone()).unwrap())
.collect::<Vec<FileInfo>>(),
));
let dupstore = Arc::new(DashMap::new());
Processor::sizewise(
Arc::new(Params::default()),
Arc::new(AtomicBool::new(true)),
dupstore.clone(),
file_queue,
Arc::new(MultiProgress::new()),
)?;
assert_eq!(dupstore.len(), 2);
Ok(())
let hashed = super::hash_candidates(vec![info], false, 300, false, &cache);
assert_eq!(hashed[0].0, 0xDEAD_BEEF);
}
#[test]
fn sizewise_sorting_two_files_of_same_size() -> Result<()> {
let root = TempDir::new()?;
let files = [
(root.path().join("fileone.bin"), generate_bytes(282624)),
(root.path().join("filetwo.bin"), generate_bytes(282624)),
];
fn hash_candidates_recomputes_when_mtime_differs() {
let root = TempDir::new().unwrap();
let path = root.path().join("f.bin");
let mut f = File::create_new(&path).unwrap();
f.write_all(b"real content for cache miss test").unwrap();
let info = FileInfo::new(path.clone()).unwrap();
let mtime = crate::cache::mtime_nanos(info.modified);
let real = info.initpages_hash(300).unwrap();
for (fpath, content) in files.iter() {
let mut f = File::create_new(fpath)?;
f.write_all(content)?;
}
let mut cache = crate::cache::Cache::disabled();
cache.record(&path, info.size, mtime + 1, false, 0xDEAD_BEEF, 0);
let file_queue = Arc::new(Mutex::new(
files
.iter()
.map(|f| FileInfo::new(f.0.clone()).unwrap())
.collect::<Vec<FileInfo>>(),
));
let dupstore = Arc::new(DashMap::new());
Processor::sizewise(
Arc::new(Params::default()),
Arc::new(AtomicBool::new(true)),
dupstore.clone(),
file_queue,
Arc::new(MultiProgress::new()),
)?;
assert_eq!(dupstore.len(), 1);
Ok(())
let hashed = super::hash_candidates(vec![info], false, 300, false, &cache);
assert_eq!(hashed[0].0, real);
assert_ne!(hashed[0].0, 0xDEAD_BEEF);
}
}

283
src/resolver.rs Normal file
View File

@@ -0,0 +1,283 @@
use std::cmp::Ordering;
use unicode_segmentation::UnicodeSegmentation;
use anyhow::{bail, Result};
use bytesize::ByteSize;
use crate::fileinfo::FileInfo;
use crate::formatter::Formatter;
use crate::params::Params;
use crate::pipeline::DedupReport;
#[derive(Debug, Clone, Copy, PartialEq, Eq, clap::ValueEnum)]
pub enum KeepStrategy {
Newest,
Oldest,
First,
Last,
Shortest,
Shallowest,
}
fn path_string(file: &FileInfo) -> String {
file.path.to_string_lossy().into_owned()
}
fn path_len(file: &FileInfo) -> usize {
file.path.to_string_lossy().graphemes(true).count()
}
fn depth(file: &FileInfo) -> usize {
file.path.iter().count()
}
fn compare_keep(a: &FileInfo, b: &FileInfo, strategy: KeepStrategy) -> Ordering {
match strategy {
KeepStrategy::Newest => b
.modified
.cmp(&a.modified)
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::Oldest => a
.modified
.cmp(&b.modified)
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::First => path_string(a).cmp(&path_string(b)),
KeepStrategy::Last => path_string(b).cmp(&path_string(a)),
KeepStrategy::Shortest => path_len(a)
.cmp(&path_len(b))
.then_with(|| path_string(a).cmp(&path_string(b))),
KeepStrategy::Shallowest => depth(a)
.cmp(&depth(b))
.then_with(|| path_len(a).cmp(&path_len(b)))
.then_with(|| path_string(a).cmp(&path_string(b))),
}
}
pub fn select_keeper(files: &[FileInfo], strategy: KeepStrategy) -> usize {
files
.iter()
.enumerate()
.min_by(|(_, a), (_, b)| compare_keep(a, b, strategy))
.map(|(index, _)| index)
.unwrap_or(0)
}
pub fn run(
report: &DedupReport,
strategy: KeepStrategy,
force: bool,
params: &Params,
) -> Result<()> {
if report.groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
}
let group_count = report.groups.len();
let mut victim_count: u64 = 0;
let mut victim_bytes: u64 = 0;
let mut freed_bytes: u64 = 0;
let mut failures: u64 = 0;
for group in &report.groups {
let keeper = select_keeper(&group.files, strategy);
if !force {
println!(
"KEEP {}",
Formatter::human_path(&group.files[keeper], params, report.max_path_len)
.unwrap_or_default()
);
}
for (index, file) in group.files.iter().enumerate() {
if index == keeper {
continue;
}
victim_count += 1;
victim_bytes += file.size;
match force {
false => println!(
"DELETE {} {}",
Formatter::human_path(file, params, report.max_path_len).unwrap_or_default(),
Formatter::human_filesize(file).unwrap_or_default()
),
true => match std::fs::remove_file(&file.path) {
Ok(_) => {
freed_bytes += file.size;
println!("deleted {}", file.path.display());
}
Err(_) => {
failures += 1;
println!("FAILED {}", file.path.display());
}
},
}
}
}
match force {
false => println!(
"\n{group_count} groups, would free {}. Re-run with --force to delete.",
ByteSize::b(victim_bytes)
),
true => println!(
"\ndeleted {} files, freed {} across {group_count} groups.",
victim_count - failures,
ByteSize::b(freed_bytes)
),
}
if failures > 0 {
bail!("{failures} deletion(s) failed");
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use std::path::PathBuf;
use std::time::{Duration, SystemTime};
use crate::params::Params;
use crate::pipeline::{DedupReport, DuplicateGroup};
use std::fs::File;
use std::io::Write;
use tempfile::TempDir;
fn file(path: &str, mtime_secs: u64) -> FileInfo {
FileInfo {
path: PathBuf::from(path).into_boxed_path(),
size: 0,
modified: SystemTime::UNIX_EPOCH + Duration::from_secs(mtime_secs),
}
}
#[test]
fn newest_keeps_greatest_mtime() {
let files = vec![file("/a/x", 10), file("/a/y", 30), file("/a/z", 20)];
assert_eq!(select_keeper(&files, KeepStrategy::Newest), 1);
}
#[test]
fn oldest_keeps_least_mtime() {
let files = vec![file("/a/x", 10), file("/a/y", 30), file("/a/z", 20)];
assert_eq!(select_keeper(&files, KeepStrategy::Oldest), 0);
}
#[test]
fn first_keeps_lexicographically_smallest_path() {
let files = vec![file("/a/y", 10), file("/a/x", 10), file("/a/z", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::First), 1);
}
#[test]
fn last_keeps_lexicographically_greatest_path() {
let files = vec![file("/a/y", 10), file("/a/x", 10), file("/a/z", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Last), 2);
}
#[test]
fn shortest_keeps_fewest_chars() {
let files = vec![file("/aaa/bbb", 10), file("/a/b", 10), file("/aa/bb", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Shortest), 1);
}
#[test]
fn shallowest_keeps_fewest_components() {
let files = vec![file("/a/b/c/d", 10), file("/a/b", 10), file("/a/b/c", 10)];
assert_eq!(select_keeper(&files, KeepStrategy::Shallowest), 1);
}
#[test]
fn newest_tiebreak_is_smallest_path_regardless_of_order() {
let forward = vec![file("/a/y", 30), file("/a/x", 30)];
let reversed = vec![file("/a/x", 30), file("/a/y", 30)];
assert_eq!(
forward[select_keeper(&forward, KeepStrategy::Newest)].path,
reversed[select_keeper(&reversed, KeepStrategy::Newest)].path
);
assert_eq!(
forward[select_keeper(&forward, KeepStrategy::Newest)]
.path
.to_string_lossy(),
"/a/x"
);
}
fn write_dup_report(root: &TempDir, names: &[&str]) -> DedupReport {
let files = names
.iter()
.map(|name| {
let path = root.path().join(name);
let mut f = File::create_new(&path).unwrap();
f.write_all(b"identical duplicate payload").unwrap();
FileInfo::new(path).unwrap()
})
.collect::<Vec<_>>();
DedupReport {
groups: vec![DuplicateGroup { hash: 0, files }],
max_path_len: 0,
}
}
#[test]
fn force_deletes_victims_and_keeps_the_keeper() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
super::run(&report, KeepStrategy::First, true, &params).unwrap();
assert!(root.path().join("a.bin").exists());
assert!(!root.path().join("b.bin").exists());
}
#[test]
fn dry_run_deletes_nothing() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
super::run(&report, KeepStrategy::First, false, &params).unwrap();
assert!(root.path().join("a.bin").exists());
assert!(root.path().join("b.bin").exists());
}
#[test]
fn force_reports_error_when_a_deletion_fails() {
let root = TempDir::new().unwrap();
let report = write_dup_report(&root, &["a.bin", "b.bin"]);
let params = Params {
dir: Some(root.path().into()),
..Default::default()
};
std::fs::remove_file(root.path().join("b.bin")).unwrap();
let result = super::run(&report, KeepStrategy::First, true, &params);
assert!(result.is_err());
}
#[test]
fn empty_report_is_ok() {
let params = Params::default();
let report = DedupReport {
groups: vec![],
max_path_len: 0,
};
assert!(super::run(&report, KeepStrategy::First, false, &params).is_ok());
}
}

View File

@@ -1,26 +1,25 @@
use crate::{fileinfo::FileInfo, params::Params};
use crate::{fileinfo::FileInfo, params::Params, pipeline::spinner};
use anyhow::Result;
use indicatif::{MultiProgress, ProgressBar, ProgressStyle};
use std::sync::{Arc, Mutex};
use std::{path::Path, time::Duration};
use std::path::Path;
use globwalk::{GlobWalker, GlobWalkerBuilder};
pub struct Scanner {
pub directory: Box<Path>,
pub filetypes: Option<String>,
pub min_depth: Option<usize>,
pub max_depth: Option<usize>,
pub include_types: Option<String>,
pub exclude_types: Option<String>,
pub min_size: Option<u64>,
pub follow_links: bool,
pub progress: bool,
}
impl Scanner {
pub fn new(app_args: Arc<Params>) -> Result<Self> {
pub fn new(app_args: &Params) -> Result<Self> {
Ok(Self {
directory: app_args.get_directory()?.into_boxed_path(),
filetypes: app_args.get_types(),
include_types: app_args.types.clone(),
exclude_types: app_args.exclude_types.clone(),
min_depth: app_args.min_depth,
max_depth: app_args.max_depth,
min_size: app_args.get_min_size(),
@@ -29,11 +28,21 @@ impl Scanner {
})
}
fn scan_patterns(&self) -> Result<String> {
Ok(match &self.filetypes {
Some(ftypes) => format!("**/*{{{ftypes}}}"),
None => "**/*".to_string(),
})
fn scan_patterns(&self) -> Result<Vec<String>> {
let include_types = match &self.include_types {
Some(ftypes) => Some(format!("**/*.{{{ftypes}}}")),
None => Some("**/*".to_string()),
};
let exclude_types = self
.exclude_types
.as_ref()
.map(|ftypes| format!("!**/*.{{{ftypes}}}"));
Ok(vec![include_types, exclude_types]
.into_iter()
.flatten()
.collect())
}
fn attach_link_opts(&self, walker: GlobWalkerBuilder) -> Result<GlobWalkerBuilder> {
@@ -53,10 +62,11 @@ impl Scanner {
None => Ok(walker),
}
}
fn build_walker(&self) -> Result<GlobWalker> {
let walker = Ok(GlobWalkerBuilder::from_patterns(
self.directory.clone(),
&[self.scan_patterns()?],
&self.scan_patterns()?,
))
.and_then(|walker| self.attach_walker_min_depth(walker))
.and_then(|walker| self.attach_walker_max_depth(walker))
@@ -65,36 +75,142 @@ impl Scanner {
Ok(walker.build()?)
}
pub fn scan(
&self,
files: Arc<Mutex<Vec<FileInfo>>>,
progress_bar_box: Arc<MultiProgress>,
) -> Result<()> {
let progress_bar = match self.progress {
true => progress_bar_box.add(ProgressBar::new_spinner()),
false => ProgressBar::hidden(),
};
pub fn scan(&self) -> Result<Vec<FileInfo>> {
let bar = spinner(self.progress, "paths mapped");
let min_size = self.min_size.unwrap_or(0);
let progress_style = ProgressStyle::with_template("[{elapsed_precise}] {pos:>7} {msg}")?;
progress_bar.set_style(progress_style);
progress_bar.enable_steady_tick(Duration::from_millis(50));
progress_bar.set_message("paths mapped");
let min_size = self.min_size.unwrap_or_default();
self.build_walker()?
let files: Vec<FileInfo> = self
.build_walker()?
.filter_map(Result::ok)
.map(|entity| entity.into_path())
.inspect(|_path| progress_bar.inc(1))
.inspect(|_path| bar.inc(1))
.filter(|path| path.is_file())
.map(FileInfo::new)
.filter_map(Result::ok)
.filter(|file| file.size > min_size)
.for_each(|file| {
let mut flock = files.lock().unwrap();
flock.push(file);
});
.filter(|file| file.size >= min_size)
.collect();
progress_bar.finish_with_message("paths mapped");
Ok(())
bar.finish_with_message("paths mapped");
Ok(files)
}
}
#[cfg(test)]
mod tests {
use crate::params::Params;
use std::fs::File;
use tempfile::TempDir;
use super::Scanner;
#[test]
fn ensure_file_include_type_filter_includes_expected_file_types() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
types: Some(String::from("js,csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-css-file.css").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
}
#[test]
fn ensure_file_exclude_type_filter_excludes_expected_file_types() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
exclude_types: Some(String::from("js,csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-css-file.css").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
}
#[test]
fn complex_file_type_params() {
let root =
TempDir::with_prefix("deduplicator_test_root").expect("unable to create tempdir");
[
"this-is-a-js-file.js",
"this-is-a-css-file.css",
"this-is-a-csv-file.csv",
"this-is-a-rust-file.rs",
]
.iter()
.for_each(|path| {
File::create_new(root.path().join(path)).unwrap_or_else(|_| {
panic!("unable to create file {path}");
});
});
let params = Params {
types: Some(String::from("js,csv,rs")),
exclude_types: Some(String::from("csv")),
dir: Some(root.path().into()),
..Default::default()
};
let scanner = Scanner::new(&params).expect("scanner initialization failed");
let files = scanner.scan().expect("scanning failed.");
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-js-file.js").to_str().unwrap()));
assert!(files.iter().all(|f| f.path.to_str().unwrap()
!= root.path().join("this-is-a-csv-file.csv").to_str().unwrap()));
assert!(files.iter().any(|f| f.path.to_str().unwrap()
== root.path().join("this-is-a-rust-file.rs").to_str().unwrap()));
}
}

View File

@@ -1,109 +0,0 @@
use std::sync::atomic::{AtomicBool, AtomicU64};
use std::sync::{Arc, Mutex};
use crate::processor::Processor;
use crate::scanner::Scanner;
use anyhow::Result;
use dashmap::DashMap;
use indicatif::{MultiProgress, ProgressDrawTarget};
use rand::Rng;
use threadpool::ThreadPool;
use crate::fileinfo::FileInfo;
use crate::params::Params;
pub struct Server {
filequeue: Arc<Mutex<Vec<FileInfo>>>,
sw_duplicate_set: Arc<DashMap<u64, Vec<FileInfo>>>,
pub hw_duplicate_set: Arc<DashMap<u128, Vec<FileInfo>>>,
threadpool: ThreadPool,
app_args: Arc<Params>,
pub max_file_path_len: Arc<AtomicU64>,
}
impl Server {
pub fn new(opts: Params) -> Self {
Self {
filequeue: Arc::new(Mutex::new(Vec::new())),
sw_duplicate_set: Arc::new(DashMap::new()),
hw_duplicate_set: Arc::new(DashMap::new()),
threadpool: ThreadPool::new(4),
app_args: Arc::new(opts),
max_file_path_len: Arc::new(AtomicU64::new(0)),
}
}
pub fn start(&self) -> Result<()> {
let progbarbox = Arc::new(MultiProgress::new());
let mut rng = rand::rng();
let seed: i64 = rng.random();
if !self.app_args.progress {
progbarbox.set_draw_target(ProgressDrawTarget::hidden());
}
let app_args_clone_for_sc = self.app_args.clone();
let app_args_clone_for_sw = self.app_args.clone();
let app_args_clone_for_hw = self.app_args.clone();
let file_queue_clone_sc = self.filequeue.clone();
let file_queue_clone_pr = self.filequeue.clone();
let scanner_finished = Arc::new(AtomicBool::new(false));
let sw_sort_finished = Arc::new(AtomicBool::new(false));
let sfin_sc_tr_cl = scanner_finished.clone();
let sfin_pr_tr_cl = scanner_finished.clone();
let swfin_pr_tr_sw = sw_sort_finished.clone();
let swfin_pr_tr_hw = sw_sort_finished.clone();
let store_dupl_sw_for_sw = self.sw_duplicate_set.clone();
let store_dupl_sw_for_hw = self.sw_duplicate_set.clone();
let store_dupl_hw = self.hw_duplicate_set.clone();
let max_file_path_len_clone = self.max_file_path_len.clone();
let progbarbox_sc_clone = progbarbox.clone();
let progbarbox_pr_clone_for_sw = progbarbox.clone();
let progbarbox_pr_clone_for_hw = progbarbox.clone();
self.threadpool.execute(move || {
Scanner::new(app_args_clone_for_sc)
.expect("unable to initialize scanner.")
.scan(file_queue_clone_sc, progbarbox_sc_clone)
.expect("scanner failed.");
sfin_sc_tr_cl.store(true, std::sync::atomic::Ordering::Relaxed);
});
self.threadpool.execute(move || {
Processor::sizewise(
app_args_clone_for_sw,
sfin_pr_tr_cl,
store_dupl_sw_for_sw,
file_queue_clone_pr,
progbarbox_pr_clone_for_sw,
)
.expect("sizewise scanner failed.");
swfin_pr_tr_sw.store(true, std::sync::atomic::Ordering::Relaxed);
});
self.threadpool.execute(move || {
Processor::hashwise(
app_args_clone_for_hw,
store_dupl_sw_for_hw,
store_dupl_hw,
progbarbox_pr_clone_for_hw,
max_file_path_len_clone,
seed,
swfin_pr_tr_hw,
)
.expect("sizewise scanner failed.");
});
progbarbox.clear()?;
self.threadpool.join();
Ok(())
}
}

499
src/tui/app.rs Normal file
View File

@@ -0,0 +1,499 @@
use std::path::PathBuf;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use crate::resolver::{select_keeper, KeepStrategy};
pub const STRATEGIES: [KeepStrategy; 6] = [
KeepStrategy::Newest,
KeepStrategy::Oldest,
KeepStrategy::First,
KeepStrategy::Last,
KeepStrategy::Shortest,
KeepStrategy::Shallowest,
];
pub fn strategy_label(strategy: KeepStrategy) -> &'static str {
match strategy {
KeepStrategy::Newest => "newest",
KeepStrategy::Oldest => "oldest",
KeepStrategy::First => "first",
KeepStrategy::Last => "last",
KeepStrategy::Shortest => "shortest",
KeepStrategy::Shallowest => "shallowest",
}
}
#[derive(Clone, Copy, PartialEq)]
pub enum Focus {
Groups,
Files,
}
#[derive(Clone, Copy, PartialEq)]
pub enum StrategyScope {
CurrentGroup,
AllGroups,
}
#[derive(Clone, Copy, PartialEq)]
pub enum Popup {
None,
Strategy { scope: StrategyScope },
ConfirmDelete,
}
pub enum Key {
Up,
Down,
Tab,
Space,
Enter,
Esc,
Char(char),
}
pub enum Outcome {
Continue,
Quit,
Delete,
Open(PathBuf),
}
pub struct Group {
pub files: Vec<FileInfo>,
pub marked: Vec<bool>,
}
impl Group {
pub fn size(&self) -> u64 {
self.files.first().map(|f| f.size).unwrap_or(0)
}
}
pub struct App {
pub groups: Vec<Group>,
pub group_cursor: usize,
pub file_cursor: usize,
pub focus: Focus,
pub popup: Popup,
pub strategy_cursor: usize,
pub status: Option<String>,
pub should_quit: bool,
}
fn clamp_add(cur: usize, delta: isize, max: usize) -> usize {
(cur as isize + delta).clamp(0, max as isize) as usize
}
impl App {
pub fn from_groups(groups: Vec<DuplicateGroup>) -> Self {
let groups = groups
.into_iter()
.map(|g| Group {
marked: vec![false; g.files.len()],
files: g.files,
})
.collect();
Self {
groups,
group_cursor: 0,
file_cursor: 0,
focus: Focus::Files,
popup: Popup::None,
strategy_cursor: 0,
status: None,
should_quit: false,
}
}
pub fn move_group(&mut self, delta: isize) {
if self.groups.is_empty() {
return;
}
self.group_cursor = clamp_add(self.group_cursor, delta, self.groups.len() - 1);
let len = self.groups[self.group_cursor].files.len();
self.file_cursor = self.file_cursor.min(len.saturating_sub(1));
}
pub fn move_file(&mut self, delta: isize) {
if let Some(g) = self.groups.get(self.group_cursor) {
self.file_cursor = clamp_add(self.file_cursor, delta, g.files.len().saturating_sub(1));
}
}
pub fn toggle_mark(&mut self) {
let fc = self.file_cursor;
if let Some(g) = self.groups.get_mut(self.group_cursor) {
if fc >= g.marked.len() {
return;
}
if !g.marked[fc] && g.marked.iter().filter(|m| !**m).count() <= 1 {
return;
}
g.marked[fc] = !g.marked[fc];
}
}
pub fn apply_strategy(&mut self, strategy: KeepStrategy, scope: StrategyScope) {
let indices: Vec<usize> = match scope {
StrategyScope::CurrentGroup => vec![self.group_cursor],
StrategyScope::AllGroups => (0..self.groups.len()).collect(),
};
for gi in indices {
if let Some(g) = self.groups.get_mut(gi) {
if g.files.len() < 2 {
continue;
}
let keeper = select_keeper(&g.files, strategy);
for i in 0..g.marked.len() {
g.marked[i] = i != keeper;
}
}
}
}
pub fn marked_count(&self) -> usize {
self.groups
.iter()
.flat_map(|g| g.marked.iter())
.filter(|m| **m)
.count()
}
pub fn reclaimable_bytes(&self) -> u64 {
self.groups
.iter()
.flat_map(|g| g.files.iter().zip(&g.marked))
.filter(|(_, m)| **m)
.map(|(f, _)| f.size)
.sum()
}
pub fn marked_paths(&self) -> Vec<PathBuf> {
self.groups
.iter()
.flat_map(|g| g.files.iter().zip(&g.marked))
.filter(|(_, m)| **m)
.map(|(f, _)| f.path.to_path_buf())
.collect()
}
pub fn apply_deletion(&mut self, deleted: &[PathBuf]) {
use std::collections::HashSet;
let removed: HashSet<&std::path::Path> = deleted.iter().map(|p| p.as_path()).collect();
for g in &mut self.groups {
let mut files = Vec::new();
let mut marked = Vec::new();
for (i, f) in g.files.iter().enumerate() {
if !removed.contains(&*f.path) {
files.push(f.clone());
marked.push(g.marked[i]);
}
}
g.files = files;
g.marked = marked;
}
self.groups.retain(|g| g.files.len() >= 2);
if self.groups.is_empty() {
self.should_quit = true;
self.group_cursor = 0;
self.file_cursor = 0;
return;
}
self.group_cursor = self.group_cursor.min(self.groups.len() - 1);
let len = self.groups[self.group_cursor].files.len();
self.file_cursor = self.file_cursor.min(len.saturating_sub(1));
}
pub fn handle_key(&mut self, key: Key) -> Outcome {
match self.popup {
Popup::Strategy { scope } => self.handle_strategy_key(key, scope),
Popup::ConfirmDelete => self.handle_confirm_key(key),
Popup::None => self.handle_main_key(key),
}
}
fn handle_main_key(&mut self, key: Key) -> Outcome {
self.status = None;
match key {
Key::Char('q') | Key::Esc => {
self.should_quit = true;
Outcome::Quit
}
Key::Tab => {
self.focus = match self.focus {
Focus::Groups => Focus::Files,
Focus::Files => Focus::Groups,
};
Outcome::Continue
}
Key::Up | Key::Char('k') => {
self.move_in_focus(-1);
Outcome::Continue
}
Key::Down | Key::Char('j') => {
self.move_in_focus(1);
Outcome::Continue
}
Key::Space => {
if self.focus == Focus::Files {
self.toggle_mark();
}
Outcome::Continue
}
Key::Char('s') => {
self.open_strategy(StrategyScope::CurrentGroup);
Outcome::Continue
}
Key::Char('S') => {
self.open_strategy(StrategyScope::AllGroups);
Outcome::Continue
}
Key::Char('o') => self.open_current(),
Key::Char('d') => {
if self.marked_count() > 0 {
self.popup = Popup::ConfirmDelete;
} else {
self.status = Some("nothing marked".to_string());
}
Outcome::Continue
}
_ => Outcome::Continue,
}
}
fn handle_strategy_key(&mut self, key: Key, scope: StrategyScope) -> Outcome {
match key {
Key::Up | Key::Char('k') => {
self.strategy_cursor = self.strategy_cursor.saturating_sub(1);
}
Key::Down | Key::Char('j') => {
self.strategy_cursor = (self.strategy_cursor + 1).min(STRATEGIES.len() - 1);
}
Key::Enter => {
let strategy = STRATEGIES[self.strategy_cursor];
self.apply_strategy(strategy, scope);
self.popup = Popup::None;
}
Key::Esc => self.popup = Popup::None,
_ => {}
}
Outcome::Continue
}
fn handle_confirm_key(&mut self, key: Key) -> Outcome {
match key {
Key::Char('y') => {
self.popup = Popup::None;
Outcome::Delete
}
Key::Char('n') | Key::Esc => {
self.popup = Popup::None;
Outcome::Continue
}
_ => Outcome::Continue,
}
}
fn move_in_focus(&mut self, delta: isize) {
match self.focus {
Focus::Groups => self.move_group(delta),
Focus::Files => self.move_file(delta),
}
}
fn open_strategy(&mut self, scope: StrategyScope) {
self.strategy_cursor = 0;
self.popup = Popup::Strategy { scope };
}
fn open_current(&self) -> Outcome {
match self.groups.get(self.group_cursor).and_then(|g| g.files.get(self.file_cursor)) {
Some(f) => Outcome::Open(f.path.to_path_buf()),
None => Outcome::Continue,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use std::path::PathBuf;
use std::time::{Duration, UNIX_EPOCH};
fn file(path: &str, size: u64, mtime_secs: u64) -> FileInfo {
FileInfo {
path: PathBuf::from(path).into_boxed_path(),
size,
modified: UNIX_EPOCH + Duration::from_secs(mtime_secs),
}
}
fn group(files: Vec<FileInfo>) -> DuplicateGroup {
DuplicateGroup { hash: 0, files }
}
fn sample() -> App {
App::from_groups(vec![
group(vec![file("/a/x", 10, 100), file("/a/y", 10, 200)]),
group(vec![file("/b/p", 5, 50), file("/b/q", 5, 60), file("/b/r", 5, 70)]),
])
}
#[test]
fn move_group_clamps_at_bounds() {
let mut app = sample();
app.move_group(-1);
assert_eq!(app.group_cursor, 0);
app.move_group(1);
assert_eq!(app.group_cursor, 1);
app.move_group(5);
assert_eq!(app.group_cursor, 1);
}
#[test]
fn move_group_reclamps_file_cursor() {
let mut app = sample();
app.group_cursor = 1;
app.file_cursor = 2;
app.move_group(-1);
assert_eq!(app.group_cursor, 0);
assert_eq!(app.file_cursor, 1);
}
#[test]
fn move_file_clamps() {
let mut app = sample();
app.move_file(-1);
assert_eq!(app.file_cursor, 0);
app.move_file(10);
assert_eq!(app.file_cursor, 1);
}
#[test]
fn toggle_mark_blocks_marking_the_last_survivor() {
let mut app = sample();
app.group_cursor = 0;
app.file_cursor = 0;
app.toggle_mark();
assert!(app.groups[0].marked[0]);
app.file_cursor = 1;
app.toggle_mark();
assert!(!app.groups[0].marked[1]);
}
#[test]
fn toggle_mark_allows_unmark() {
let mut app = sample();
app.file_cursor = 0;
app.toggle_mark();
app.toggle_mark();
assert!(!app.groups[0].marked[0]);
}
#[test]
fn apply_strategy_current_group_marks_all_but_keeper() {
let mut app = sample();
app.group_cursor = 0;
app.apply_strategy(KeepStrategy::Newest, StrategyScope::CurrentGroup);
assert_eq!(app.groups[0].marked, vec![true, false]);
assert_eq!(app.groups[1].marked, vec![false, false, false]);
}
#[test]
fn apply_strategy_all_groups() {
let mut app = sample();
app.apply_strategy(KeepStrategy::Oldest, StrategyScope::AllGroups);
assert_eq!(app.groups[0].marked, vec![false, true]);
assert_eq!(app.groups[1].marked, vec![false, true, true]);
}
#[test]
fn totals_sum_marked_only() {
let mut app = sample();
app.apply_strategy(KeepStrategy::Newest, StrategyScope::AllGroups);
assert_eq!(app.marked_count(), 3);
assert_eq!(app.reclaimable_bytes(), 20);
}
#[test]
fn marked_paths_lists_marked_files() {
let mut app = sample();
app.group_cursor = 0;
app.file_cursor = 0;
app.toggle_mark();
assert_eq!(app.marked_paths(), vec![PathBuf::from("/a/x")]);
}
#[test]
fn apply_deletion_removes_files_and_drops_small_groups() {
let mut app = sample();
app.apply_deletion(&[PathBuf::from("/a/x")]);
assert_eq!(app.groups.len(), 1);
assert_eq!(app.groups[0].files.len(), 3);
}
#[test]
fn apply_deletion_emptying_everything_sets_quit() {
let mut app = App::from_groups(vec![group(vec![file("/a/x", 1, 1), file("/a/y", 1, 2)])]);
app.apply_deletion(&[PathBuf::from("/a/x")]);
assert!(app.groups.is_empty());
assert!(app.should_quit);
}
#[test]
fn handle_key_quits_on_q() {
let mut app = sample();
assert!(matches!(app.handle_key(Key::Char('q')), Outcome::Quit));
assert!(app.should_quit);
}
#[test]
fn handle_key_space_marks_in_files_focus() {
let mut app = sample();
app.focus = Focus::Files;
app.handle_key(Key::Space);
assert!(app.groups[0].marked[0]);
}
#[test]
fn handle_key_strategy_popup_apply_flow() {
let mut app = sample();
app.handle_key(Key::Char('S'));
assert!(matches!(app.popup, Popup::Strategy { .. }));
app.handle_key(Key::Enter);
assert!(matches!(app.popup, Popup::None));
assert_eq!(app.groups[0].marked, vec![true, false]);
}
#[test]
fn handle_key_delete_requires_marks_then_confirms() {
let mut app = sample();
app.handle_key(Key::Char('d'));
assert!(matches!(app.popup, Popup::None));
app.focus = Focus::Files;
app.handle_key(Key::Space);
app.handle_key(Key::Char('d'));
assert!(matches!(app.popup, Popup::ConfirmDelete));
assert!(matches!(app.handle_key(Key::Char('y')), Outcome::Delete));
assert!(matches!(app.popup, Popup::None));
}
#[test]
fn handle_key_open_returns_path() {
let mut app = sample();
app.focus = Focus::Files;
match app.handle_key(Key::Char('o')) {
Outcome::Open(p) => assert_eq!(p, PathBuf::from("/a/x")),
_ => panic!("expected Open"),
}
}
}

129
src/tui/mod.rs Normal file
View File

@@ -0,0 +1,129 @@
pub mod app;
mod ui;
use std::io::{self, Stdout};
use anyhow::Result;
use ratatui::backend::CrosstermBackend;
use ratatui::crossterm::event::{self, Event, KeyCode, KeyEvent, KeyEventKind};
use ratatui::crossterm::execute;
use ratatui::crossterm::terminal::{
disable_raw_mode, enable_raw_mode, EnterAlternateScreen, LeaveAlternateScreen,
};
use ratatui::Terminal;
use crate::params::Params;
use crate::pipeline::DedupReport;
use app::{App, Key, Outcome};
type Tui = Terminal<CrosstermBackend<Stdout>>;
pub fn run(report: DedupReport, params: &Params) -> Result<()> {
if report.groups.is_empty() {
println!("No duplicates found matching your search criteria.");
return Ok(());
}
let mut app = App::from_groups(report.groups);
let mut terminal = setup_terminal()?;
let result = event_loop(&mut terminal, &mut app, params);
restore_terminal(&mut terminal)?;
result
}
fn setup_terminal() -> Result<Tui> {
install_panic_hook();
enable_raw_mode()?;
let mut stdout = io::stdout();
execute!(stdout, EnterAlternateScreen)?;
Ok(Terminal::new(CrosstermBackend::new(stdout))?)
}
fn restore_terminal(terminal: &mut Tui) -> Result<()> {
let _ = disable_raw_mode();
let _ = execute!(terminal.backend_mut(), LeaveAlternateScreen);
let _ = terminal.show_cursor();
Ok(())
}
fn install_panic_hook() {
let original = std::panic::take_hook();
std::panic::set_hook(Box::new(move |info| {
let _ = disable_raw_mode();
let _ = execute!(io::stdout(), LeaveAlternateScreen);
original(info);
}));
}
fn event_loop(terminal: &mut Tui, app: &mut App, params: &Params) -> Result<()> {
loop {
terminal.draw(|frame| ui::draw(frame, app, params))?;
if app.should_quit {
break;
}
if let Event::Key(key) = event::read()? {
if key.kind != KeyEventKind::Press {
continue;
}
if let Some(translated) = translate(key) {
match app.handle_key(translated) {
Outcome::Quit => {}
Outcome::Continue => {}
Outcome::Delete => perform_deletion(app),
Outcome::Open(path) => {
if let Err(err) = open::that(&path) {
app.status = Some(format!("open failed: {err}"));
}
}
}
}
}
if app.should_quit {
break;
}
}
Ok(())
}
fn translate(key: KeyEvent) -> Option<Key> {
Some(match key.code {
KeyCode::Up => Key::Up,
KeyCode::Down => Key::Down,
KeyCode::Tab => Key::Tab,
KeyCode::Enter => Key::Enter,
KeyCode::Esc => Key::Esc,
KeyCode::Char(' ') => Key::Space,
KeyCode::Char(c) => Key::Char(c),
_ => return None,
})
}
fn perform_deletion(app: &mut App) {
let paths = app.marked_paths();
let mut deleted = Vec::new();
let mut freed: u64 = 0;
let mut failures: u64 = 0;
for path in &paths {
let size = std::fs::metadata(path).map(|m| m.len()).unwrap_or(0);
match std::fs::remove_file(path) {
Ok(_) => {
deleted.push(path.clone());
freed += size;
}
Err(_) => failures += 1,
}
}
app.apply_deletion(&deleted);
app.status = Some(match failures {
0 => format!("deleted {}, freed {}", deleted.len(), bytesize::ByteSize::b(freed)),
_ => format!(
"deleted {}, {failures} failed, freed {}",
deleted.len(),
bytesize::ByteSize::b(freed)
),
});
}

211
src/tui/ui.rs Normal file
View File

@@ -0,0 +1,211 @@
use ratatui::layout::{Constraint, Direction, Layout, Rect};
use ratatui::style::Style;
use ratatui::text::Line;
use ratatui::widgets::{Block, Borders, Clear, List, ListItem, ListState, Paragraph};
use ratatui::Frame;
use crate::params::Params;
use crate::formatter::Formatter;
use super::app::{strategy_label, App, Focus, Popup, StrategyScope, STRATEGIES};
pub fn draw(frame: &mut Frame, app: &App, params: &Params) {
let rows = Layout::default()
.direction(Direction::Vertical)
.constraints([Constraint::Min(1), Constraint::Length(1)])
.split(frame.area());
let panes = Layout::default()
.direction(Direction::Horizontal)
.constraints([Constraint::Percentage(35), Constraint::Percentage(65)])
.split(rows[0]);
draw_groups(frame, app, panes[0]);
draw_files(frame, app, params, panes[1]);
draw_footer(frame, app, rows[1]);
match app.popup {
Popup::Strategy { scope } => draw_strategy_popup(frame, app, scope),
Popup::ConfirmDelete => draw_confirm_popup(frame, app),
Popup::None => {}
}
}
fn pane_block(title: &str, focused: bool) -> Block<'_> {
let block = Block::default().borders(Borders::ALL).title(title.to_string());
match focused {
true => block.border_style(Style::new().yellow()),
false => block,
}
}
fn draw_groups(frame: &mut Frame, app: &App, area: Rect) {
let items: Vec<ListItem> = app
.groups
.iter()
.enumerate()
.map(|(i, g)| {
ListItem::new(format!(
"{:>3} {} files {}",
i + 1,
g.files.len(),
bytesize::ByteSize::b(g.size())
))
})
.collect();
let list = List::new(items)
.block(pane_block("Groups", app.focus == Focus::Groups))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.group_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_files(frame: &mut Frame, app: &App, params: &Params, area: Rect) {
let title = format!(
"Group {}/{}",
(app.group_cursor + 1).min(app.groups.len().max(1)),
app.groups.len()
);
let items: Vec<ListItem> = match app.groups.get(app.group_cursor) {
Some(g) => g
.files
.iter()
.enumerate()
.map(|(i, f)| {
let box_char = if g.marked[i] { "[x]" } else { "[ ]" };
let path = Formatter::human_path(f, params, 0).unwrap_or_default();
let size = Formatter::human_filesize(f).unwrap_or_default();
let mtime = Formatter::human_mtime(f).unwrap_or_default();
let line = format!("{box_char} {} {} {}", path.trim_end(), size.trim_start(), mtime);
match g.marked[i] {
true => ListItem::new(Line::from(line).style(Style::new().red())),
false => ListItem::new(line),
}
})
.collect(),
None => Vec::new(),
};
let list = List::new(items)
.block(pane_block(&title, app.focus == Focus::Files))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.file_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_footer(frame: &mut Frame, app: &App, area: Rect) {
let hints = "[Tab]pane [Space]mark [s]group [S]all [o]pen [d]elete [q]uit";
let left = match &app.status {
Some(s) => s.clone(),
None => format!(
"marked {} · reclaim {}",
app.marked_count(),
bytesize::ByteSize::b(app.reclaimable_bytes())
),
};
let text = format!("{left} {hints}");
frame.render_widget(Paragraph::new(text), area);
}
fn centered_rect(width: u16, height: u16, area: Rect) -> Rect {
let x = area.x + area.width.saturating_sub(width) / 2;
let y = area.y + area.height.saturating_sub(height) / 2;
Rect {
x,
y,
width: width.min(area.width),
height: height.min(area.height),
}
}
fn draw_strategy_popup(frame: &mut Frame, app: &App, scope: StrategyScope) {
let title = match scope {
StrategyScope::CurrentGroup => "Keep in this group",
StrategyScope::AllGroups => "Keep in ALL groups",
};
let items: Vec<ListItem> = STRATEGIES
.iter()
.map(|s| ListItem::new(strategy_label(*s)))
.collect();
let area = centered_rect(28, (STRATEGIES.len() as u16) + 2, frame.area());
frame.render_widget(Clear, area);
let list = List::new(items)
.block(Block::default().borders(Borders::ALL).title(title))
.highlight_style(Style::new().reversed());
let mut state = ListState::default();
state.select(Some(app.strategy_cursor));
frame.render_stateful_widget(list, area, &mut state);
}
fn draw_confirm_popup(frame: &mut Frame, app: &App) {
let area = centered_rect(48, 3, frame.area());
frame.render_widget(Clear, area);
let text = format!(
"Delete {} files, reclaim {}? [y/N]",
app.marked_count(),
bytesize::ByteSize::b(app.reclaimable_bytes())
);
frame.render_widget(
Paragraph::new(text).block(Block::default().borders(Borders::ALL).title("Confirm")),
area,
);
}
#[cfg(test)]
mod tests {
use super::*;
use crate::fileinfo::FileInfo;
use crate::pipeline::DuplicateGroup;
use crate::tui::app::App;
use ratatui::backend::TestBackend;
use ratatui::Terminal;
use std::path::PathBuf;
use std::time::{Duration, UNIX_EPOCH};
fn sample_app() -> App {
let files = vec![
FileInfo {
path: PathBuf::from("/tmp/report.pdf").into_boxed_path(),
size: 4_000_000,
modified: UNIX_EPOCH + Duration::from_secs(1_700_000_000),
},
FileInfo {
path: PathBuf::from("/tmp/report-copy.pdf").into_boxed_path(),
size: 4_000_000,
modified: UNIX_EPOCH + Duration::from_secs(1_700_100_000),
},
];
App::from_groups(vec![DuplicateGroup { hash: 0, files }])
}
#[test]
fn draw_renders_without_panic_and_shows_content() {
let app = sample_app();
let params = crate::params::Params::default();
let backend = TestBackend::new(100, 24);
let mut terminal = Terminal::new(backend).unwrap();
terminal.draw(|f| draw(f, &app, &params)).unwrap();
let text: String = terminal
.backend()
.buffer()
.content()
.iter()
.map(|cell| cell.symbol())
.collect();
assert!(text.contains("Groups"));
assert!(text.contains("report"));
assert!(text.contains("marked"));
}
}