Files
J621/ROADMAP.md
T
JakeBreath 09405d1a0f Match local files to e621: MD5 lookups, manual links, batch scans
- MediaItem gains e621_match_status (unknown/matched/not_found/deleted)
  and e621_checked_at, backfilled for existing matched items.
- Server-side e621 client (apps/library/e621.py) using the user's stored
  credentials, throttled to 2 req/s, with typed errors.
- Matching service: MD5 lookup, manual post linking (flags MD5
  mismatches), unlink, metadata refresh, deleted-post detection.
- Detail actions POST /api/files/J-x/match/ and /unlink/ (uploader or
  staff only).
- Background library scans: MatchTask + /api/matches/ with missing/all
  scopes, progress polling, cancel and stale-task reaping; the scan
  counts toward the footer's Active Workers. Same pass available as
  manage.py match_e621 for cron.
- Library gains not_found/deleted status filters; the detail page adds
  an e621 match card (check / link by post ID / unlink) and the metadata
  card warns when a post was deleted on e621.
2026-09-17 13:41:08 -05:00

5.2 KiB

J621 Roadmap / To-Dos

Working list of what's still missing, roughly in priority order. Check items off as they land.

1. Library management

  • Duplicates engine
    • Exact MD5 duplicate detection with grouped results
    • Perceptual similarity (aHash, dHash, pHash, wHash) with threshold slider and algorithm toggles (imagehash server-side)
    • Visual similarity groups: pagination, selection, delete and dismiss
  • Delete & storage page
    • Storage overview (watched folder, media folder, temp)
    • Delete by J-ID with preview grid and bulk selection (plus per-copy deletion of duplicate locations)
    • Temp folder cleanup
  • Library search upgrades
    • Search by tags (custom + e621 tags) with a Filename / Tags / Both selector
    • Tag cloud in the sidebar (click to search, hidden from guests for blacklisted items)
    • Status filter (matched / custom / unknown — not_found/deleted arrive with the e621 match cache)

2. e621 integration

  • Download progress bar on the detail view
    • Backend download task (status, progress %, downloaded/total bytes, speed)
    • Progress endpoint the SPA polls; cancel support
    • Frontend progress bar with speed and cancel on /detail/<post>
    • The status footer's "Active Workers" now counts running download tasks
  • Download to client — streams the e621 original straight to the browser (works for guests; host-restricted proxy, no library write)
  • Match local files to e621
    • Match by MD5 from the library detail, plus manual post ID linking (with an MD5-mismatch warning) and unlink
    • Batch cache status for the whole library (matched / not_found / deleted) via background scan tasks (/api/matches/) and manage.py match_e621
    • Show match status + e621 metadata in the library: status badge and match card on the detail, and not_found / deleted status filters
  • IQDB reverse search
    • Search from a local file (POST /iqdb_queries.json)
    • Show candidate posts to link
  • Follows
    • Follow tags/pools, Followed screens with unseen badges
    • "Followed" nav entry
  • F shortcut to favorite a post, tag finder in the Ctrl+K palette

3. Uploads & ingestion pipeline

Files now stage first and are resolved before entering the library.

  • Backend staging storage
    • TempUpload model: user, md5, filename, temp path, status (pending / visual_match / completed / error), resolution
    • Files land in a temp folder first; only completed uploads move into the watched library folder
    • Pending/upload list endpoint, file serving, discard endpoint
    • cleanup_temp_uploads command for old staged files
  • Auto-upload / auto-match
    • MD5 computed on staging; exact duplicates resolve immediately
    • MD5 batch-checked against e621; matches auto-complete with post metadata stored and rating seeded
  • IQDB similarity on upload (SPA-driven)
    • Automatic + manual IQDB checks with candidate posts
    • "Visual Similarity Detected" state with candidate picker
    • Perceptual-hash comparison against the library (staged uploads are flagged with their library matches as soon as they land)
  • Upload UI
    • Three-column board: Pending & Unmatched / Visual Similarity Detected / Auto-uploaded & Indexed
    • Metadata modal (link to e621 post, IQDB candidates, custom metadata)
    • Per-file progress plus batch processing indicator

4. Staff tools

  • Optimization modal (image quality, GIF/APNG resolution+FPS+compression, video bitrate/presets/two-pass/hw-accel) with original vs processed preview and override
  • Stats dashboard
    • CPU / RAM / GPU / disk usage
    • Active jobs list (this is what makes the footer "Active Workers" real)
    • Live log tail

5. Shell & polish

  • Toasts instead of inline messages / confirm dialogs
  • Mobile drawer polish for metadata panels (design spec §layout)
  • Profile pictures: staff Users page sets avatars from library J-IDs (self-service picker in Account still pending)
  • Profile extras (per-user browse preferences)
  • Command palette: tag finder, recent searches

6. Infrastructure

  • Guest blacklist refresh on a timer (refresh_guest_blacklist via cron/systemd)
  • Production setup: build the SPA, serve via Nginx (static + /media + /library), systemd unit for Waitress
  • Automated tests (backend API + frontend components)
  • Backfill e621 metadata for items downloaded before metadata was stored (re-download or a "match" action)

Dependencies / notes

  • Duplicates and upload visual-similarity share the perceptual hashing layer (imagehash/imgdd server-side + a hash cache table).
  • Download progress, optimization jobs and the stats "Active Workers" count all want the same background-task/progress primitive — design it once.
  • IQDB and e621 matching depend on e621 credentials being configured; the SPA talks to e621 directly for browsing, while the backend e621 client (apps/library/e621.py) handles metadata matching and batch scans.