Files
J621/ROADMAP.md
T
JakeBreath 8ebda6ab20 Docker deployment (2 images, 3 composes, Tailscale funnels) + committed security suite
Test suite (the manual audit harness, now a real test):
- backend/apps/core/tests/test_security.py: 20 transactional tests across
  guest visibility, object ownership, staged-upload/similarity privacy,
  staff role boundaries, deletion rules, encrypted credentials, throttling
  and the download allowlist. Uses temp media folders and clears cache.
  Needs a one-time GRANT on test_j621 (documented in the module + README).

Images (deploy/J621-Frontend, deploy/J621-Backend, repo root as context):
- Frontend: node build -> static nginx with SPA fallback, asset caching and
  an internal health endpoint.
- Backend: gunicorn + whitenoise (admin static collected at build), ffmpeg
  for video thumbnails, migrations applied on start, GIT_HASH build arg so
  the shell's version pill shows the commit.

Composes (distinct project names so they coexist with the dev stack):
- compose.yml (both), compose.frontend.yml, compose.backend.yml.
- A shared nginx proxy service is the only entry point (no host nginx, no
  published host ports): /api,/admin,/static,/health -> backend, everything
  else -> SPA; both upstreams resolve at request time so one config serves
  all variants.
- A Tailscale sidecar per compose shares the nginx network namespace;
  serve.default/frontend/backend.json use funnel ports 443 and 8443 only
  (10000 is the remaining allowance) with ${TS_CERT_DOMAIN} substitution.

Registry: push_frontend.sh / push_backend.sh / push_all.sh build multi-arch
images and push :latest + :<sha> to the Gitea registry, following the
existing Packs-site pattern.

Docs: deploy/README.md + .env.example, ROADMAP section 6 updated,
AGENTS.md deployment and test notes.

Verified: all three composes validate, both nginx configs pass nginx -t,
both images build, the combined stack boots against real MariaDB/Redis
(migrations applied, /health ok, whitenoise serving admin static, env=prod,
git hash baked in), the proxy serves the SPA and routes /api, and
manage.py test apps.core.tests passes 20/20.
2026-09-18 00:51:29 -05:00

9.9 KiB

J621 Roadmap / To-Dos

Working list of what's still missing, roughly in priority order. Check items off as they land.

1. Library management

  • Duplicates engine

    • Exact MD5 duplicate detection with grouped results
    • Perceptual similarity (aHash, dHash, pHash, wHash) with threshold slider and algorithm toggles (imagehash server-side)
    • Visual similarity groups: pagination, selection, delete and dismiss
  • Delete & storage page

    • Storage overview (watched folder, media folder, temp)
    • Delete by J-ID with preview grid and bulk selection (plus per-copy deletion of duplicate locations)
    • Temp folder cleanup
  • Library search upgrades

    • Search by tags (custom + e621 tags) with a Filename / Tags / Both selector
    • Tag cloud in the sidebar (click to search, hidden from guests for blacklisted items)
    • Status filter (matched / not_found / deleted / custom / unknown)
  • Ephemeral similarity check (/similar)

    • Drop a file: exact MD5 match, perceptual matches against the library, and e621 IQDB candidates (auto-run for images)
    • Nothing enters the library: temp files are wiped on startup, after SIMILARITY_TTL_MINUTES (default 30), on demand and by manage.py cleanup_similarity

2. e621 integration

  • Download progress bar on the detail view
    • Backend download task (status, progress %, downloaded/total bytes, speed)
    • Progress endpoint the SPA polls; cancel support
    • Frontend progress bar with speed and cancel on /detail/<post>
    • The status footer's "Active Workers" now counts running download tasks
  • Download to client — streams the e621 original straight to the browser (works for guests; host-restricted proxy, no library write)
  • Match local files to e621
    • Match by MD5 from the library detail, plus manual post ID linking (with an MD5-mismatch warning) and unlink
    • Batch cache status for the whole library (matched / not_found / deleted) via background scan tasks (/api/matches/) and manage.py match_e621
    • Show match status + e621 metadata in the library: status badge and match card on the detail, and not_found / deleted status filters
  • IQDB reverse search
    • Search from a local file (POST /iqdb_queries.json from the library detail, using the signed raw file)
    • Show candidate posts to link (thumbnail, rating, score, favs, tags; select then confirm — same flow as uploads, exact MD5 marked)
  • Follows
    • Follow tags from the + on tag chips (online + library detail) and pools from the Follow button on a pool page (validated against e621, first feed fetched immediately)
    • Followed screens: cover cards (latest image), unseen badges and Mark seen per follow
    • Merged newest-first feed with per-follow filter and unseen-only toggle
    • Blacklisted-tag cloud built in a daemon thread, then polled (10 min TTL; the user's e621 blacklist, guest default as fallback)
    • "Followed" nav entry
    • Periodic sync commands: sync_followed_tags and sync_followed_pools (one e621 search per unique tag/pool)
  • Pools browser (/pools + /pools/<id>)
    • Index search by name; category, active/deleted and sort filters; pagination according to the e621 OpenAPI spec
    • Cover thumbnails from each pool's first post (one batched post call, blacklist-aware) and deleted-pool markers
    • Pool detail: DText description, posts kept in pool order with chunked loading and in-library badges, blacklist reveal toggle, Follow button
  • F shortcut to favorite a post (online detail), tag finder in the Ctrl+K palette (e621 tag suggestions with follow toggles, recent searches, and "search e621 for …")

3. Uploads & ingestion pipeline

Files now stage first and are resolved before entering the library.

  • Backend staging storage
    • TempUpload model: user, md5, filename, temp path, status (pending / visual_match / completed / error), resolution
    • Files land in a temp folder first; only completed uploads move into the watched library folder
    • Pending/upload list endpoint, file serving, discard endpoint
    • cleanup_temp_uploads command for old staged files
  • Auto-upload / auto-match
    • MD5 computed on staging; exact duplicates resolve immediately
    • MD5 batch-checked against e621; matches auto-complete with post metadata stored and rating seeded
  • IQDB similarity on upload (SPA-driven)
    • Automatic + manual IQDB checks with candidate posts
    • "Visual Similarity Detected" state with candidate picker
    • Perceptual-hash comparison against the library (staged uploads are flagged with their library matches as soon as they land)
  • Upload UI
    • Three-column board: Pending & Unmatched / Visual Similarity Detected / Auto-uploaded & Indexed
    • Metadata modal (link to e621 post, IQDB candidates, custom metadata)
    • Per-file progress plus batch processing indicator

4. Staff tools

  • Optimization modal (client-side, no server processing)
    • Browser pipeline in a Web Worker: Mediabunny/WebCodecs for video (hardware-accelerated where available), jSquash (MozJPEG/OxiPNG/libwebp) for images, gifenc/gifuct-js + UPNG for GIF/APNG
    • Options per file type: image quality/resolution; animation scale/fps/palette; video quality/resolution/codec/container/hardware preference (codec support detected per browser)
    • Original vs processed previews with sizes and savings, progress bar with ETA and a "Processing the file…" state
    • Apply overwrites the same J-ID: POST /api/files/J-x/optimize/ replaces the file(s) and recomputes MD5/size/perceptual hashes (409 when the result matches another item, 400 when it is identical)
    • Animated formats are detected by header (APNG acTL chunk, WebP VP8X flag), so APNGs keep their frames and animated WebP is refused with a notice instead of being flattened (browsers have no animated WebP encoder)
  • Stats dashboard (/stats, staff)
    • CPU / RAM / GPU / disk usage (psutil + nvidia-smi; capacity bars follow the design system's 80% / 95% colour thresholds)
    • Active jobs list (downloads + match scans with progress) plus the recently finished ones
    • Live log tail with level highlighting and an auto-scroll toggle (root logging now also writes a rotating backend/logs/j621.log)

5. Shell & polish

  • 18+ entry screen: an age check gates the app before anything renders (remembered per browser in localStorage)
  • Toasts for action results (success/error) plus a shared confirm dialog for destructive actions; form-field validation stays inline
  • Mobile bottom sheets for metadata panels (<768px): detail asides (Library/Online/Similar) and long pool descriptions
  • Profile pictures: staff Users page and a self-service Account picker (searchable library grid, remove supported) set avatars from library J-IDs
  • Profile extras: per-user landing page, default rating filter/sort, items per page and thumbnail size (synced to the account, applied on load)
  • Backend-less local mode: with no backend connected the SPA runs on the e621-facing pages only (Online, Pools) using credentials stored in the browser, and the shell offers a "Setup Backend" button instead
  • Command palette: tag finder, recent searches, navigation commands (Ctrl+K / ⌘K, reachable from anywhere in the shell)

6. Infrastructure

  • Guest blacklist refresh on a timer (in-container scheduler)
  • Follow sync on a timer (in-container scheduler, e.g. every 30 minutes)
  • Similarity temp cleanup on a timer (in-container scheduler; the TTL also cleans lazily when new checks are created)
  • Production setup: two Docker images (SPA on static nginx, API on gunicorn + whitenoise with ffmpeg) and three compose variants (both / frontend-only / backend-only) behind a shared nginx proxy service, each with a Tailscale sidecar and a funnel serve config (ports 443 / 8443 / 10000 only). Multi-arch images are pushed to the Gitea registry by deploy/push_*.sh
  • Security/permission test suite (manage.py test apps.core.tests): guest visibility, object ownership, upload/similarity privacy, role boundaries, deletion rules, encrypted credentials, throttling, SSRF allowlist
  • Broader automated tests (more backend API coverage + frontend components)
  • Backfill e621 metadata for items downloaded before metadata was stored — covered by the match scan (/api/matches/ scope all, or per-item Recheck) and manage.py match_e621 --scope all

Dependencies / notes

  • Duplicates and upload visual-similarity share the perceptual hashing layer (imagehash server-side; no external hashing service).
  • Download progress, optimization and the stats "Active Workers" count share the same task style; optimization itself runs in the browser (the home server is too weak for ffmpeg) and only the result is uploaded.
  • IQDB and e621 matching depend on e621 credentials being configured; the SPA talks to e621 directly for browsing, while the backend e621 client (apps/library/e621.py) handles metadata matching and batch scans.
  • Anything touching e621 endpoints follows the OpenAPI spec (https://e621.wiki/openapi.yaml) — see AGENTS.md for how to fetch and which response shapes to watch out for.
  • Security review done (Aug 2026): permissions verified per role with a repeatable harness, e621 keys encrypted at rest, rate limits added and the download paths host-allowlisted. See AGENTS.md for the rules to keep.