Files
J621/ROADMAP.md
T
JakeBreath 1adb761c8d Fix following past 48 entries and make the e621 page size configurable
Follow lists were paginated at the API default of 48, but the SPA treats
them as complete sets: the tag/pool toggles read their state from page one
(so the 49th follow looked unfollowed and its spinner waited for a page that
could never contain it) and the Followed page rendered only 48 cards while
showing that as the count. Both follow endpoints are now unpaginated — they
are per-user sets and still restricted to the caller's rows — and the three
consumers take plain arrays.

Post visibility: the old J621-Django online view fetched limit=320 (e621's
maximum) while ours hard-coded 48, and fetchPostsByIds capped id batches at
100. The Online browser now has a 'Posts per page' setting (48/100/200/320)
in its sidebar, mirrored in Account -> Browsing preferences, stored per user
as e621_per_page and also used for pool loading; the id-batch cap is raised
to 320.

Tests: follow list shape/isolation (4) and preference validation/merge (3)
added; the full backend suite is 46 green. Live-checked the array response
shape and the preference bounds (200 accepted, 500 rejected).
2026-09-18 19:06:18 -05:00

197 lines
11 KiB
Markdown

# J621 Roadmap / To-Dos
Working list of what's still missing, roughly in priority order. Check items off as
they land.
## 1. Library management
- [x] **Duplicates engine**
- [x] Exact MD5 duplicate detection with grouped results
- [x] Perceptual similarity (aHash, dHash, pHash, wHash) with threshold slider
and algorithm toggles (`imagehash` server-side)
- [x] Visual similarity groups: pagination, selection, delete and dismiss
- [x] **Delete & storage page**
- [x] Storage overview (watched folder, media folder, temp)
- [x] Delete by J-ID with preview grid and bulk selection (plus per-copy
deletion of duplicate locations)
- [x] Temp folder cleanup
- [x] **Library search upgrades**
- [x] Search by tags (custom + e621 tags) with a Filename / Tags / Both selector
- [x] Tag cloud in the sidebar (click to search, hidden from guests for
blacklisted items)
- [x] Status filter (matched / not_found / deleted / custom / unknown)
- [x] **Random image** endpoint and page: `GET /api/random/` (aliases
`/random`, `/random/`) with `rating=s,q,e` filters; fastfetch mode
(`?fastfetch=1` or a Fastfetch User-Agent) only returns png/jpg/gif as
JSON with a signed absolute link, for the fish_greeting scripts; the
`/random` SPA page rolls with rating pills and the `R` key
- **Scoped `j621r_…` greeting tokens** (Account → Shell tokens, `/tokens`):
only `/api/random/` accepts them, they are stored hashed, shown once and
revocable; `install.fish` in `extras/fish_greeting` fills them into the
shell greeting config
- [x] **Ephemeral similarity check** (`/similar`)
- [x] Drop a file: exact MD5 match, perceptual matches against the library,
and e621 IQDB candidates (auto-run for images)
- [x] Nothing enters the library: temp files are wiped on startup, after
`SIMILARITY_TTL_MINUTES` (default 30), on demand and by
`manage.py cleanup_similarity`
## 2. e621 integration
- [x] **Download progress bar on the detail view**
- [x] Backend download task (status, progress %, downloaded/total bytes, speed)
- [x] Progress endpoint the SPA polls; cancel support
- [x] Frontend progress bar with speed and cancel on `/detail/<post>`
- [x] The status footer's "Active Workers" now counts running download tasks
- [x] **Download to client** — streams the e621 original straight to the
browser (works for guests; host-restricted proxy, no library write)
- [x] **Match local files to e621**
- [x] Match by MD5 from the library detail, plus manual post ID linking
(with an MD5-mismatch warning) and unlink
- [x] Batch cache status for the whole library (`matched` / `not_found` /
`deleted`) via background scan tasks (`/api/matches/`) and
`manage.py match_e621`
- [x] Show match status + e621 metadata in the library: status badge and
match card on the detail, and `not_found` / `deleted` status filters
- [x] **IQDB reverse search**
- [x] Search from a local file (`POST /iqdb_queries.json` from the library detail,
using the signed raw file)
- [x] Show candidate posts to link (thumbnail, rating, score, favs, tags;
select then confirm — same flow as uploads, exact MD5 marked)
- [x] **Follows**
- [x] Follow tags from the `+` on tag chips (online + library detail) and
pools from the Follow button on a pool page (validated against e621,
first feed fetched immediately)
- [x] Followed screens: cover cards (latest image), unseen badges and
Mark seen per follow
- [x] Merged newest-first feed with per-follow filter and unseen-only toggle
- [x] Blacklisted-tag cloud built in a daemon thread, then polled
(10 min TTL; the user's e621 blacklist, guest default as fallback)
- [x] "Followed" nav entry
- [x] Periodic sync commands: `sync_followed_tags` and
`sync_followed_pools` (one e621 search per unique tag/pool)
- [x] **Pools browser** (`/pools` + `/pools/<id>`)
- [x] Index search by name; category, active/deleted and sort filters;
pagination according to the e621 OpenAPI spec
- [x] Cover thumbnails from each pool's first post (one batched post call,
blacklist-aware) and deleted-pool markers
- [x] Pool detail: DText description, posts kept in pool order with chunked
loading and in-library badges, blacklist reveal toggle, Follow button
- [x] `F` shortcut to favorite a post (online detail), tag finder in the Ctrl+K
palette (e621 tag suggestions with follow toggles, recent searches, and
"search e621 for …")
## 3. Uploads & ingestion pipeline
Files now stage first and are resolved before entering the library.
- [x] **Backend staging storage**
- [x] `TempUpload` model: user, md5, filename, temp path, status
(`pending` / `visual_match` / `completed` / `error`), resolution
- [x] Files land in a temp folder first; only completed uploads move into
the watched library folder
- [x] Pending/upload list endpoint, file serving, discard endpoint
- [x] `cleanup_temp_uploads` command for old staged files
- [x] **Auto-upload / auto-match**
- [x] MD5 computed on staging; exact duplicates resolve immediately
- [x] MD5 batch-checked against e621; matches auto-complete with post
metadata stored and rating seeded
- [x] **IQDB similarity on upload** (SPA-driven)
- [x] Automatic + manual IQDB checks with candidate posts
- [x] "Visual Similarity Detected" state with candidate picker
- [x] Perceptual-hash comparison against the library (staged uploads are
flagged with their library matches as soon as they land)
- [x] **Upload UI**
- [x] Three-column board: Pending & Unmatched / Visual Similarity Detected /
Auto-uploaded & Indexed
- [x] Metadata modal (link to e621 post, IQDB candidates, custom metadata)
- [x] Per-file progress plus batch processing indicator
## 4. Staff tools
- [x] **Optimization modal** (client-side, no server processing)
- [x] Browser pipeline in a Web Worker: Mediabunny/WebCodecs for video
(hardware-accelerated where available), jSquash (MozJPEG/OxiPNG/libwebp)
for images, gifenc/gifuct-js + UPNG for GIF/APNG
- [x] Options per file type: image quality/resolution; animation
scale/fps/palette; video quality/resolution/codec/container/hardware
preference (codec support detected per browser)
- [x] Original vs processed previews with sizes and savings, progress bar
with ETA and a "Processing the file…" state
- [x] Apply overwrites the same J-ID: `POST /api/files/J-x/optimize/` replaces
the file(s) and recomputes MD5/size/perceptual hashes (409 when the
result matches another item, 400 when it is identical)
- [x] Animated formats are detected by header (APNG `acTL` chunk, WebP
`VP8X` flag), so APNGs keep their frames and animated WebP is refused
with a notice instead of being flattened (browsers have no animated
WebP encoder)
- [x] **Stats dashboard** (`/stats`, staff)
- [x] CPU / RAM / GPU / disk usage (`psutil` + `nvidia-smi`; capacity bars
follow the design system's 80% / 95% colour thresholds)
- [x] Active jobs list (downloads + match scans with progress) plus the
recently finished ones
- [x] Live log tail with level highlighting and an auto-scroll toggle
(root logging now also writes a rotating `backend/logs/j621.log`)
## 5. Shell & polish
- [x] 18+ entry screen: an age check gates the app before anything renders
(remembered per browser in `localStorage`)
- [x] Toasts for action results (success/error) plus a shared confirm dialog for
destructive actions; form-field validation stays inline
- [x] Mobile bottom sheets for metadata panels (<768px): detail asides
(Library/Online/Similar) and long pool descriptions
- [x] Profile pictures: staff Users page and a self-service Account picker
(searchable library grid, remove supported) set avatars from library J-IDs
- [x] Profile extras: per-user landing page, default rating filter/sort, items
per page and thumbnail size (synced to the account, applied on load),
plus e621 posts per page (48/100/200/320) for the Online browser and
pool loading
- [x] Backend-less local mode: with no backend connected the SPA runs on the
e621-facing pages only (Online, Pools) using credentials stored in the
browser, and the shell offers a "Setup Backend" button instead
- [x] Command palette: tag finder, recent searches, navigation commands
(Ctrl+K / ⌘K, reachable from anywhere in the shell)
## 6. Infrastructure
- [x] Guest blacklist refresh on a timer (the composes' `scheduler` service;
`J621_BLACKLIST_EVERY`, default daily)
- [x] Follow sync on a timer (same scheduler: `sync_followed_tags` +
`sync_followed_pools`, default every 30 minutes)
- [x] Similarity temp cleanup on a timer (same scheduler, default hourly;
the TTL also cleans lazily when new checks are created)
- [x] Production setup: two Docker images (SPA on static nginx, API on
gunicorn + whitenoise with ffmpeg) and three compose variants (both /
frontend-only / backend-only) behind a shared nginx proxy service, each
with a Tailscale sidecar and a funnel serve config (ports 443 / 8443 /
10000 only). Multi-arch images are pushed to the Gitea registry by
`deploy/push_*.sh`
- [x] Security/permission test suite (`manage.py test apps.core.tests`):
guest visibility, object ownership, upload/similarity privacy, role
boundaries, deletion rules, encrypted credentials, throttling, SSRF
allowlist
- [ ] Broader automated tests (more backend API coverage + frontend
components)
- [x] Backfill e621 metadata for items downloaded before metadata was stored —
covered by the match scan (`/api/matches/` scope `all`, or per-item
Recheck) and `manage.py match_e621 --scope all`
## Dependencies / notes
- Duplicates and upload visual-similarity share the perceptual hashing layer
(`imagehash` server-side; no external hashing service).
- Download progress, optimization and the stats "Active Workers" count share
the same task style; optimization itself runs in the browser (the home
server is too weak for ffmpeg) and only the result is uploaded.
- IQDB and e621 matching depend on e621 credentials being configured; the SPA
talks to e621 directly for browsing, while the backend e621 client
(`apps/library/e621.py`) handles metadata matching and batch scans.
- Anything touching e621 endpoints follows the OpenAPI spec
(https://e621.wiki/openapi.yaml) — see AGENTS.md for how to fetch and
which response shapes to watch out for.
- Security review done (Aug 2026): permissions verified per role with a
repeatable harness, e621 keys encrypted at rest, rate limits added and the
download paths host-allowlisted. See AGENTS.md for the rules to keep.