Signed media URLs embedded the current second (TimestampSigner), so every
API response re-minted every raw/thumbnail/staged URL and the browser
re-downloaded each file on every poll or navigation. Responses also carried
no cache headers at all.
- sign with a plain Signer plus a bucket-quantized exp (7d TTL, 24h bucket),
so a URL is byte-identical across responses and rotates once a day; legacy
TimestampSigner URLs stay accepted for one release
- add a v=<md5> version parameter to library media URLs so replacing a file
under the same J-ID (the optimize flow) busts caches exactly when needed
- serve_file now sends ETag/Last-Modified and a private Cache-Control and
answers conditional requests with 304; library media gets max-age 6d +
immutable, staged/similarity files 1h
- build cached 480px JPEG thumbnails for images (Pillow, keyed by MD5 under
MEDIA_ROOT/thumbs) instead of serving full-size originals through the
thumbnail endpoint; the library grid uses thumbnail_url for images too
Batching must not start while files are still being uploaded, and the MD5
phase must move a whole chunk at once instead of one resolve per file:
- the upload queue drains completely first (failed uploads included) before
any matching starts;
- phase 1 asks e621 for every md5 (75 per posts.json request), builds the
md5 -> post map from the response, and sends the matches to the new
POST /api/uploads/link-bulk/ action, so a whole 75-file chunk moves into
Indexed in a single board update;
- link-bulk indexes the staged file directly when the post's MD5 matches
(identical bytes), so there is no per-file download round trip;
- phase 2 runs local visual similarity for whatever stayed pending, phase 3
the IQDB queue.
Verified end to end with real e621 files: one md5 query for the batch, one
link-bulk call, both matching files flipped to Indexed together, then the
visual and IQDB phases. 23 library tests green (link-bulk, visual phase,
deferred visual matching).
Uploads were doing md5 + local visual matching inside the upload request
(backend create) while the frontend later ran its own e621 MD5 pass, so the
pipeline looked interleaved per file. Now every step is a phase applied to
the whole batch in order:
1. upload (fast: md5 + exact-duplicate check only),
2. e621 MD5 lookup, 75 md5: metatags per posts.json request,
3. local visual similarity, one file at a time via the new
POST /api/uploads/<id>/visual-match/ action,
4. IQDB through the existing serial queue.
The board shows the active phase with its own progress bar (e621 MD5 in
peach, visual in lavender, IQDB in teal) and every step updates the staged
list as it lands. Verified from a headless run: one batched posts.json
request for 10 files, then 10 visual-match calls, then IQDB.
The modal held a snapshot of the staged upload, so IQDB results that landed
from the background check queue never appeared until it was closed and
reopened — the only hint a check was running was the e621 request history.
It now follows the live uploads query, so candidates, progress and errors
show up in place.
Related gaps fixed along the way:
- files flagged by the local visual-similarity check were skipped by the
IQDB pass entirely (only 'pending' files were checked), so their modal
could only ever show 'already in your library'; unresolved files of both
statuses are now checked, and the check button shows on visual-match
cards too;
- a check with no candidates posted nothing, leaving 'never checked' and
'checked, no match' indistinguishable; results are stored even when
empty and the modal now says which one it is;
- per-file failures surface in the modal instead of being swallowed, the
modal shows a spinner while the query runs and a check now/re-check
button, and auto-runs skip files already checked (and videos, since IQDB
is image-only).
Backend production code unchanged; tests pin the empty-result recording
(18 library tests, full suite 56 green).
Upload board:
- the tile grid no longer re-sorts itself as files finish (that reshuffled
the list under the cursor); it keeps insertion order, uses auto-fill tiles
of ~150px so they hold a readable size, scrolls inside a 60vh area and no
longer chains the page scroll (overscroll-contain);
- the files currently in flight are pinned in a small live strip above the
grid (name, percent, bar) so progress stays visible while the grid is
scrolled with hundreds of tiles.
Bulk rating: a 'bulk rate' button in the Pending & Unmatched header opens a
large modal with Safe/Questionable/Explicit pills, a tickable thumbnail grid
(Select all / Clear) and one confirm that moves every selected upload into
the library with that rating. Backed by POST /api/uploads/resolve-bulk/
(temp_ids + rating, own rows only): each staged file is resolved as a custom
entry (keeps its staged tags/notes), and already-completed or foreign ids are
reported per entry instead of failing the whole batch. Built for the
358-file backlog.
Tests: 4 bulk-resolve tests (resolution with the rating, input validation,
foreign ids untouched, mixed completed+pending) — full backend suite 53
green. Verified live end to end: staged a file, bulk-resolved it as 'q', saw
J-96 created with that rating, then removed the item, temp row and test
token.
The upload board partitions /api/uploads/ into Pending / Visual similarity /
Auto-uploaded, but the endpoint was paginated at 48 — a 69-file batch
silently lost 21 entries, and the similarity sweep (which reads the same
list back after uploading) only ever saw the first page. The staged-upload
list is now unpaginated: it is a transient per-user set, still limited to
the caller's rows and the uploader role. The page takes a plain array.
Watching progress with dozens of files was also poor:
- the queue uploads three files at a time instead of strictly one at a time;
- the Uploads section now shows a batch bar and 'n/m uploaded · x%' next to
the count, so the overall progress never scrolls out of sight;
- entries are ordered active-first (uploading, queued, failed, done) so the
file being uploaded is always at the top of the grid;
- tiles are larger (4 columns at lg instead of 5);
- the header reads 'Uploading n/m…' and 'Checking n file(s) against IQDB…'
instead of a bare spinner.
Tests: staged-upload list unpaginated past 48, per-user, uploader-only
(3 new; full suite 49 green). Live-checked the bare-array response.
Backend: GET /api/random/ (aliases /random and /random/) returns a random
library image with:
- rating=s,q,e filtering (comma separated, default any);
- fastfetch mode (?fastfetch=1 or any User-Agent containing "fastfetch")
that only considers png/jpg/gif - what terminal viewers can show;
- JSON with j_id, filename, extension, rating, size, e621 id plus absolute
url/download_url/thumbnail_url. Authenticated callers get signed URLs so
fastfetch and image viewers can load them without headers; guests get
unsigned URLs and never receive hidden_from_guests items.
Tests: apps/library/tests/test_random.py (8 tests) covering the response
contract, guest signatures, image-only default, the fastfetch format
restriction (flag and User-Agent), rating filters, guest visibility and the
short alias.
Frontend: /random page with rating pills, R to roll, Open/Download and a
library link, plus navigation and command palette entries; needs a backend,
hidden in local mode.
nginx: /random negotiates on Accept so browsers keep getting the SPA while
scripts get the JSON (verified with the proxy and frontend containers).
Also fixes a regression from the SSRF change: the guest download proxy
still referenced the removed 'parsed' variable on its success path, so
every proxied download would have 500'd. Redirect hops are now covered by
tests with a mocked requests.get.