40 Commits
Author SHA1 Message Date
JakeBreath 7c2569522f Make thumbnail generation atomic and warm it on import
- write thumbnails to a .part file and os.replace() them, so concurrent
  requests never read a half-written JPEG
- a stale thumbnail plus a vanished source no longer raises through the
  request (getmtime on a missing file returned 500); it falls back cleanly
- ensure_thumbnail(item) warms the preview when a file is indexed, keeping
  image decoding out of the request path
2026-09-23 21:35:09 -05:00
JakeBreath a62195ffce Re-sign visual-match thumbnails on every detail fetch
Match rows stored a signed URL minted when the scan ran, so it aged out (or
used the pre-stable signing scheme) and the modal showed broken tiles even
after legacy signatures were fixed. Rows now carry item_id/j_id and the
detail serializer mints a fresh thumbnail URL per request; matches whose
item no longer exists are dropped.
2026-09-23 21:33:20 -05:00
JakeBreath c73a81a5f4 Fix 500 on legacy signed URLs
A TimestampSigner value is an HMAC over 'payload:timestamp', so a plain
Signer's HMAC check accepts it and the embedded timestamp then reached the
JSON decoder, raising JSONDecodeError (not BadSignature) and surfacing as a
500. That broke every stored visual-match thumbnail URL minted before the
stable scheme, so the J-ID match tiles never loaded on prod.

Detect the legacy shape by its extra separator and verify it with
TimestampSigner; malformed input returns None instead of raising.
2026-09-23 21:31:34 -05:00
JakeBreath 9ababb8b48 Stop storing completed uploads; announce them through a live feed
Every auto-matched, duplicate or manually resolved upload left a completed
TempUpload row on the board until it was dismissed by hand, so the rows
accumulated without bound and the bulk dismiss (capped at 1000 ids) failed
once there were more. The original app never stored these: they are
notifications, not records.

- complete_temp_upload now appends {filename, J-ID, resolution, post} to a
  bounded recent_completions feed on UploadRun and deletes the staged row
- staging duplicates never create a board record either; the create response
  carries the J-ID and preview so the SPA can show the card immediately
- status_payload returns the feed (newest first, signed thumbnails) for the
  live board; finalize_round counts deleted matches in processed
- resolve/link-bulk return synthetic completion payloads
- migration 0012 adds the field and purges the existing completed backlog
  (and any stray staged files) on deploy
2026-09-23 18:24:38 -05:00
JakeBreath 37085c5dac Make media URLs stable and cacheable, add real image thumbnails
Signed media URLs embedded the current second (TimestampSigner), so every
API response re-minted every raw/thumbnail/staged URL and the browser
re-downloaded each file on every poll or navigation. Responses also carried
no cache headers at all.

- sign with a plain Signer plus a bucket-quantized exp (7d TTL, 24h bucket),
  so a URL is byte-identical across responses and rotates once a day; legacy
  TimestampSigner URLs stay accepted for one release
- add a v=<md5> version parameter to library media URLs so replacing a file
  under the same J-ID (the optimize flow) busts caches exactly when needed
- serve_file now sends ETag/Last-Modified and a private Cache-Control and
  answers conditional requests with 304; library media gets max-age 6d +
  immutable, staged/similarity files 1h
- build cached 480px JPEG thumbnails for images (Pillow, keyed by MD5 under
  MEDIA_ROOT/thumbs) instead of serving full-size originals through the
  thumbnail endpoint; the library grid uses thumbnail_url for images too
2026-09-23 18:13:52 -05:00
JakeBreath e2697c0a78 Stop throttling signed media and ease the browser's e621 queue
Signed media URLs are fetched by <img>/<video> tags without an
Authorization header, so they were charged to the anonymous 120/min
bucket: past that, galleries and the fish-greeting download got 429 JSON
instead of image bytes. The raw/thumbnail/staged-file/similarity-file
actions are now exempt, and THROTTLE_ENABLED=false removes the general
anon+user limits for private/tailnet deployments (login/register/proxy
guards stay).

The SPA's e621 client also stops self-throttling so hard: 1s gap between
browsing calls (2.5s for the stricter IQDB endpoint) and a 15s cooldown
instead of 60s when e621 answers 429.
2026-09-22 22:32:37 -05:00
JakeBreath 474403ffe2 Upload updates 2026-09-21 09:01:01 -05:00
JakeBreath 3a07481dfc Wait for all uploads, then batch MD5 -> visual -> IQDB with bulk links
Batching must not start while files are still being uploaded, and the MD5
phase must move a whole chunk at once instead of one resolve per file:

- the upload queue drains completely first (failed uploads included) before
  any matching starts;
- phase 1 asks e621 for every md5 (75 per posts.json request), builds the
  md5 -> post map from the response, and sends the matches to the new
  POST /api/uploads/link-bulk/ action, so a whole 75-file chunk moves into
  Indexed in a single board update;
- link-bulk indexes the staged file directly when the post's MD5 matches
  (identical bytes), so there is no per-file download round trip;
- phase 2 runs local visual similarity for whatever stayed pending, phase 3
  the IQDB queue.

Verified end to end with real e621 files: one md5 query for the batch, one
link-bulk call, both matching files flipped to Indexed together, then the
visual and IQDB phases. 23 library tests green (link-bulk, visual phase,
deferred visual matching).
2026-09-19 11:10:24 -05:00
JakeBreath e2bf1c457f Make upload processing phase-based: MD5 -> visual -> IQDB over the batch
Uploads were doing md5 + local visual matching inside the upload request
(backend create) while the frontend later ran its own e621 MD5 pass, so the
pipeline looked interleaved per file. Now every step is a phase applied to
the whole batch in order:

1. upload (fast: md5 + exact-duplicate check only),
2. e621 MD5 lookup, 75 md5: metatags per posts.json request,
3. local visual similarity, one file at a time via the new
   POST /api/uploads/<id>/visual-match/ action,
4. IQDB through the existing serial queue.

The board shows the active phase with its own progress bar (e621 MD5 in
peach, visual in lavender, IQDB in teal) and every step updates the staged
list as it lands. Verified from a headless run: one batched posts.json
request for 10 files, then 10 visual-match calls, then IQDB.
2026-09-19 10:25:57 -05:00
JakeBreath 9641862515 Back off from e621 rate limits and pace requests more conservatively
e621 intermittently answers 429 to the IQDB endpoint; browsers hide that
status behind CORS ('Access-Control-Allow-Origin missing'), so the SPA
cannot read it. Treat every network-level failure as a possible rate limit
and pause all e621 traffic for a minute. The cooldown is shared through
localStorage so extra tabs respect it, requests are spaced 1.5s apart
instead of 1s, user-cancelled requests do not trigger a cooldown, and the
upload queue waits the cooldown out with a countdown instead of looking
stuck.

Server side: the per-process e621 gap goes from 0.5s to 1s so two gunicorn
workers cannot together exceed e621's 2/s hard limit.
2026-09-19 00:37:32 -05:00
JakeBreath bc7494e7be Show IQDB checks in the metadata modal and run them for visual matches
The modal held a snapshot of the staged upload, so IQDB results that landed
from the background check queue never appeared until it was closed and
reopened — the only hint a check was running was the e621 request history.
It now follows the live uploads query, so candidates, progress and errors
show up in place.

Related gaps fixed along the way:
- files flagged by the local visual-similarity check were skipped by the
  IQDB pass entirely (only 'pending' files were checked), so their modal
  could only ever show 'already in your library'; unresolved files of both
  statuses are now checked, and the check button shows on visual-match
  cards too;
- a check with no candidates posted nothing, leaving 'never checked' and
  'checked, no match' indistinguishable; results are stored even when
  empty and the modal now says which one it is;
- per-file failures surface in the modal instead of being swallowed, the
  modal shows a spinner while the query runs and a check now/re-check
  button, and auto-runs skip files already checked (and videos, since IQDB
  is image-only).

Backend production code unchanged; tests pin the empty-result recording
(18 library tests, full suite 56 green).
2026-09-18 23:42:45 -05:00
JakeBreath a761def65e Scrollable upload grid and bulk rating for the pending backlog
Upload board:
- the tile grid no longer re-sorts itself as files finish (that reshuffled
  the list under the cursor); it keeps insertion order, uses auto-fill tiles
  of ~150px so they hold a readable size, scrolls inside a 60vh area and no
  longer chains the page scroll (overscroll-contain);
- the files currently in flight are pinned in a small live strip above the
  grid (name, percent, bar) so progress stays visible while the grid is
  scrolled with hundreds of tiles.

Bulk rating: a 'bulk rate' button in the Pending & Unmatched header opens a
large modal with Safe/Questionable/Explicit pills, a tickable thumbnail grid
(Select all / Clear) and one confirm that moves every selected upload into
the library with that rating. Backed by POST /api/uploads/resolve-bulk/
(temp_ids + rating, own rows only): each staged file is resolved as a custom
entry (keeps its staged tags/notes), and already-completed or foreign ids are
reported per entry instead of failing the whole batch. Built for the
358-file backlog.

Tests: 4 bulk-resolve tests (resolution with the rating, input validation,
foreign ids untouched, mixed completed+pending) — full backend suite 53
green. Verified live end to end: staged a file, bulk-resolved it as 'q', saw
J-96 created with that rating, then removed the item, temp row and test
token.
2026-09-18 21:50:22 -05:00
JakeBreath 39307cb141 Unpaginate staged uploads and make big upload batches visible
The upload board partitions /api/uploads/ into Pending / Visual similarity /
Auto-uploaded, but the endpoint was paginated at 48 — a 69-file batch
silently lost 21 entries, and the similarity sweep (which reads the same
list back after uploading) only ever saw the first page. The staged-upload
list is now unpaginated: it is a transient per-user set, still limited to
the caller's rows and the uploader role. The page takes a plain array.

Watching progress with dozens of files was also poor:
- the queue uploads three files at a time instead of strictly one at a time;
- the Uploads section now shows a batch bar and 'n/m uploaded · x%' next to
  the count, so the overall progress never scrolls out of sight;
- entries are ordered active-first (uploading, queued, failed, done) so the
  file being uploaded is always at the top of the grid;
- tiles are larger (4 columns at lg instead of 5);
- the header reads 'Uploading n/m…' and 'Checking n file(s) against IQDB…'
  instead of a bare spinner.

Tests: staged-upload list unpaginated past 48, per-user, uploader-only
(3 new; full suite 49 green). Live-checked the bare-array response.
2026-09-18 19:58:17 -05:00
JakeBreath 770b1e5ee6 Scoped API tokens for the random endpoint, with a management page
Backend: a GreetingToken model stores only a SHA-256 hash of a j621r_…
key (shown once at creation) plus label, prefix, created/last-used. A
dedicated GreetingTokenAuthentication understands the usual
'Authorization: Token …' header but is registered only on RandomItemView
(alongside the normal token auth), so a greeting token authenticates
/api/random/ and is rejected with 401 everywhere else — exactly the scope
shell greetings need. Endpoints: GET/POST /api/auth/greeting-tokens/ and
DELETE /api/auth/greeting-tokens/{id}/ (own tokens only; the list never
returns keys or hashes).

Frontend: /tokens page (Account → Shell tokens card, command palette entry)
lists tokens with label, prefix, created/last-used and revoke (shared
confirm dialog). Creating one shows the key with Copy and 'Copy for fish'
buttons plus a pointer to extras/fish_greeting.

Tests: apps/accounts/tests/test_greeting_tokens.py — 9 tests covering
create-once semantics and hashing, hidden keys in listings, the scope
guarantee (random 200 with a signed URL; 401 on files, storage, me, tags
cloud, delete and the token list itself), unknown/revoked keys, cross-user
revocation, last-used tracking and label limits.

Verified live: created a token, rolled /random (signed URL), got 401 from
four other endpoints, saw the list omit secrets, revoked it (204) and the
same key then 401'd on /random. Full suite: 39 tests green.
2026-09-18 13:37:29 -05:00
JakeBreath f8667c1037 Add a Random image endpoint and SPA page (with fastfetch mode)
Backend: GET /api/random/ (aliases /random and /random/) returns a random
library image with:
- rating=s,q,e filtering (comma separated, default any);
- fastfetch mode (?fastfetch=1 or any User-Agent containing "fastfetch")
  that only considers png/jpg/gif - what terminal viewers can show;
- JSON with j_id, filename, extension, rating, size, e621 id plus absolute
  url/download_url/thumbnail_url. Authenticated callers get signed URLs so
  fastfetch and image viewers can load them without headers; guests get
  unsigned URLs and never receive hidden_from_guests items.

Tests: apps/library/tests/test_random.py (8 tests) covering the response
contract, guest signatures, image-only default, the fastfetch format
restriction (flag and User-Agent), rating filters, guest visibility and the
short alias.

Frontend: /random page with rating pills, R to roll, Open/Download and a
library link, plus navigation and command palette entries; needs a backend,
hidden in local mode.

nginx: /random negotiates on Accept so browsers keep getting the SPA while
scripts get the JSON (verified with the proxy and frontend containers).

Also fixes a regression from the SSRF change: the guest download proxy
still referenced the removed 'parsed' variable on its success path, so
every proxied download would have 500'd. Redirect hops are now covered by
tests with a mocked requests.get.
2026-09-18 13:06:36 -05:00
JakeBreath f86eccf9a3 Security fixes: SSRF, staff role escalation, SPA-only gating, throttling, encrypted keys
Findings from the audit (50-check harness across guest/user/uploader/staff/
admin) and their fixes:

- SSRF: 'Download to Library' and the staged-upload resolve path fetched
  any http(s) URL. services.validate_remote_url now enforces the e621
  media allowlist and open_remote re-validates every redirect hop; the
  download-task create endpoint and the guest proxy use them, so internal
  addresses (127.0.0.1, LAN, metadata) are rejected with 400.
- Privilege escalation: staff could promote users to staff and demote
  other staff. Role changes across the staff boundary now require an
  admin, matching the account-deletion rules; the Users page hides what
  the backend would refuse.
- SPA-only gating: /api/storage/ and /api/duplicates/* were readable by
  any authenticated account (absolute paths, duplicate groups) while the
  SPA only shows them to uploaders. They now require CanUpload.
- Throttling (REST_FRAMEWORK, env-overridable, counted in Redis):
  anon 120/min, user 600/min, login 5/min, register 20/hour, guest e621
  proxy 60/hour. Login now goes through a throttled view.
- e621 API keys are encrypted at rest with a Fernet key derived from
  SECRET_KEY (apps/accounts/crypto.py); a data migration encrypts existing
  rows and the column widens first. Reads decrypt transparently, legacy
  plaintext still works, and a changed SECRET_KEY reads as 'not
  configured' instead of leaking. Rotating SECRET_KEY now invalidates
  stored keys as well as signed media URLs.
- Hardening: the server refuses to start with DEBUG=False while SECRET_KEY
  is still the development default.

Verified: corrected harness 50/50 (guest visibility, IDOR, signed-URL
tamper/expiry, staged-upload/similarity privacy, role matrix, SSRF),
login throttles at the 6th attempt with 429, anon polling unaffected, the
guest proxy still reaches allowlisted hosts, live e621 auth works with the
decrypted key, and DB rows hold only ciphertext.
2026-09-18 00:21:14 -05:00
JakeBreath 99f617d296 Footer storage/backend display and a staff role that actually grants staff
Footer:
- Left is now 'Backend Storage:' with a capacity bar (blue, peach at 80%,
  red at 95% per DESIGN.md) and a used/total/free tooltip; the watched
  folder path is no longer printed. /api/status/ returns a compact storage
  summary instead of the path (the full storage page still shows paths to
  authenticated users).
- Centre shows the backend API origin (empty = same origin). Staff get a
  link to /setup to point the browser elsewhere; everyone else sees it as
  plain text. The Account 'Backend connection' card is gone — this is
  installation plumbing, not a per-user setting.
- Design spec updated to match.

Staff role:
- The custom role did nothing on several endpoints that only accepted
  Django's is_staff/is_superuser. One canonical check now exists:
  User.is_app_staff (superuser, Django staff, or the staff role), used by
  the stats/users APIs, item object permissions, can_delete, upload/
  similarity/download/match querysets, and the management commands
  (which also pick staff-role accounts for e621 sync/match and file
  ownership).

Verified with a role-only staff account (is_staff/is_superuser false):
stats/users 200, all 32 downloads + 2 scans visible, others' items
editable; the same account as role=user gets 403 for all of those.
2026-09-17 23:24:24 -05:00
JakeBreath 16907c39ca Support cross-origin frontends alongside same-origin setups
- django-cors-headers with env-driven CORS_ALLOWED_ORIGINS,
  CORS_ALLOW_ALL_ORIGINS, CORS_ALLOW_CREDENTIALS and CSRF_TRUSTED_ORIGINS;
  same-origin traffic is unaffected and a disallowed origin gets no CORS
  headers. Token auth needs no cookies, so credentials stay off by default.
- TRUST_PROXY_HEADERS=true lets a TLS-terminating proxy supply
  X-Forwarded-Proto/Host for correct absolute URLs.
- API media URLs (raw/thumbnail/upload/similarity/staged previews) are now
  absolute, built from the request host, so <img>/<video>/fetch() keep
  working when the SPA is served from another origin. Signed URLs are still
  per-user; nothing is stored in the DB.
- The SPA gains VITE_API_BASE (build-time, empty = same-origin) applied by
  a small apiUrl() helper used for XHR/fetch and the few URL fallbacks.

Verified with a throwaway instance: preflight and GET responses carry the
allowed origin, foreign origins get nothing, media GETs include CORS for
cross-origin fetch(), and payload URLs use the request host (dev :8000
unchanged).
2026-09-17 22:50:12 -05:00
JakeBreath 487d14c617 Keep the UI attached to background jobs across navigation
Jobs already run on the server — leaving the page or closing the tab does
not stop them — but the SPA lost its link to them because the task id
lived in component state. The online detail page now looks up the newest
task for the post: an active one resumes the progress bar and cancel
button, and a finished one shows "your last download for this post
finished — J-xx". The downloads list accepts a post_id filter for that
lookup.

The footer's Active Workers count is also a link to the staff stats
dashboard, which is the global view of running jobs.
2026-09-17 21:45:24 -05:00
JakeBreath 27cbfd882c Add job cancellation to the stats dashboard
- Active jobs on /stats get a cancel button wired to the existing
  download/match cancel endpoints, showing "cancelling..." and an inline
  error when the task already finished.
- Cancelling now sets the status immediately, so a task whose runner died
  in a restart stops showing as "downloading".
- Download streams use a bounded read timeout (10 s connect / 60 s read):
  a stalled socket fails within a minute (previously it could block
  forever), and a task cancelled while stalled is marked cancelled rather
  than error.
- The stats job list reaps stale download/match tasks, so phantom jobs
  never appear on the dashboard.
2026-09-17 21:41:54 -05:00
JakeBreath b024fc52d7 Add the staff stats dashboard
Backend: GET /api/stats/ (staff only) gathers psutil CPU/memory counters,
nvidia-smi GPU stats, the cached disk numbers and the running/finished
download + match jobs. Root logging now also writes a rotating file
(backend/logs/j621.log) so the dashboard can tail it, and psutil joins the
requirements. The storage payload computation is shared with the existing
storage endpoint.

Frontend: a /stats route + Stats nav entry for staff, polling every 2 s —
per-core CPU bars, memory and swap, GPUs (utilization, VRAM, temperature),
disk with the media/temp breakdown, active jobs with progress bars,
recently finished jobs with summaries, and the log tail with level colours
and an auto-scroll toggle. Section 4 of the roadmap is complete.
2026-09-17 21:36:33 -05:00
JakeBreath bab9904fc8 Client-side optimization pipeline with apply-to-J-ID
Backend:
- POST /api/files/J-x/optimize/ applies a browser-processed file: replaces
  every copy (renaming when the extension changes), recomputes MD5, size
  and perceptual hashes, seeds guest visibility; 400 when identical,
  409 when the result matches another item, owner/staff only.

Frontend (no server-side processing by design):
- optimize.worker.ts + pipelines: Mediabunny/WebCodecs for video with a
  prefer-hardware hint and per-browser codec detection; MozJPEG/OxiPNG/
  libwebp (jSquash) for images; gifuct-js+gifenc and UPNG for GIF/APNG.
- OptimizeModal: per-file-type options, original vs processed previews
  with sizes/savings, progress bar with ETA, then Apply (overwrite).
- Optimize button on the library detail for the uploader/staff.
2026-09-17 20:03:22 -05:00
JakeBreath c061d2681b Fix Download to Library crashing with a NameError
run_download_task sets e621_match_status from MediaItem, but the module
only imported DownloadTask — every download that carried e621 metadata
failed right after indexing, leaving the file in the library unlinked.
Verified the runner end-to-end with a stubbed fetch.
2026-09-17 19:28:19 -05:00
JakeBreath 1c2cb8d468 Add an ephemeral similarity check page
- /similar (nav: Similar): drop a file to get the exact MD5 match, the
  perceptual matches against the library, and e621 IQDB candidates
  (auto-run for images when credentials are configured). Read-only —
  nothing enters the library.
- SimilarityCheck model + /api/similarity/ (create/list/retrieve/delete)
  with signed preview URLs and an expires_at timestamp.
- Temp files are wiped on startup (AppConfig.ready, file-only so no
  database access during initialization), lazily past
  SIMILARITY_TTL_MINUTES (default 30, env-overridable), on delete, and
  by manage.py cleanup_similarity.
- uploadFile() takes a target path; .env.example documents the TTL.
2026-09-17 18:22:11 -05:00
JakeBreath 03dd235f3a Follows: followed tags/pools, feeds, unseen badges and blacklist cloud
Backend (new apps.follows):
- FollowedTag/FollowedPool/FollowedPost models; per-user follows with
  unseen tracking, plus FollowCloud for the cached blacklist cloud.
- Two periodic commands sharing one fetch path: sync_followed_tags and
  sync_followed_pools fetch each followed tag/pool's newest posts (one
  e621 search per unique follow), store unseen feed rows, refresh covers
  and pool metadata; both fall back to anonymous e621 access.
- API: /api/follows/tags|pools (follow, unfollow, mark seen), a merged
  feed with per-follow filtering, and /api/follows/cloud/ which rebuilds
  the blacklisted-tag cloud in a daemon thread when its 10 min cache is
  stale (polling returns building/ready).
- e621 client now supports anonymous reads; trimmed posts carry preview
  URLs for covers and feed tiles.

Frontend:
- /followed page: follow forms, cover cards with unseen badges and
  Mark seen, merged feed with filter/unseen toggle, and a blacklist
  cloud panel that polls while building. Followed nav entry added.
2026-09-17 14:09:30 -05:00
JakeBreath 09405d1a0f Match local files to e621: MD5 lookups, manual links, batch scans
- MediaItem gains e621_match_status (unknown/matched/not_found/deleted)
  and e621_checked_at, backfilled for existing matched items.
- Server-side e621 client (apps/library/e621.py) using the user's stored
  credentials, throttled to 2 req/s, with typed errors.
- Matching service: MD5 lookup, manual post linking (flags MD5
  mismatches), unlink, metadata refresh, deleted-post detection.
- Detail actions POST /api/files/J-x/match/ and /unlink/ (uploader or
  staff only).
- Background library scans: MatchTask + /api/matches/ with missing/all
  scopes, progress polling, cancel and stale-task reaping; the scan
  counts toward the footer's Active Workers. Same pass available as
  manage.py match_e621 for cron.
- Library gains not_found/deleted status filters; the detail page adds
  an e621 match card (check / link by post ID / unlink) and the metadata
  card warns when a post was deleted on e621.
2026-09-17 13:41:08 -05:00
JakeBreath db74f7ab18 Fix hidden library items not rendering in the browser
Items flagged hidden_from_guests (blacklisted tags) returned 404 for
<img> requests since tags cannot send the auth header. The API now
exposes signed raw_url/thumbnail_url fields (mirroring upload previews
and avatars), and the SPA uses them in the gallery, detail view,
duplicates and delete screens, and upload visual matches.
2026-09-17 13:25:20 -05:00
JakeBreath 4df573da43 Library search upgrades: tag search, tag cloud, status filter
Backend:
- MediaItem gains search_tags (custom + e621 tags, lowercase) and
  has_custom_data, maintained on save with a data migration backfill
- File list search accepts search_type=filename|tags|both (tag search is
  word-AND across the flattened tag text) and status=matched|custom|
  unknown filters
- New /api/tags/cloud/ endpoint (cached 2 min per guest/auth, invalidated
  on item changes and deletions) returning the most-used tags, honouring
  guest visibility

Frontend:
- Library sidebar: Filename/Tags/Both selector, status pill toggles
  (persisted), and a clickable tag cloud that runs a tag search
- Roadmap updated
2026-09-17 13:19:17 -05:00
JakeBreath cd490b0a23 Duplicates, delete & storage, users page with J-ID avatars
Backend:
- Perceptual hashes (aHash/dHash/pHash/wHash via imagehash, no imgdd)
  stored on items, computed on upload/download and by the new
  compute_visual_hashes command
- Duplicates API: exact duplicates (multi-location items), visual matches
  for one item, union-find similarity groups with pagination
- Delete API with ownership/staff checks, per-item and per-copy deletion,
  watched-folder path validation; storage overview and temp cleanup;
  file list accepts j_ids batches
- Staged uploads are flagged visual_match with their library matches
  (threshold via VISUAL_MATCH_THRESHOLD)
- Staff users API: list with upload counts, set role and avatar by J-ID;
  User.avatar FK with signed avatar URLs
- Download threads close their DB connection and stale tasks are reaped,
  keeping behaviour Gunicorn-friendly

Frontend:
- /duplicates: exact duplicate groups with per-copy delete, visual
  similarity controls, search similar to a J-ID, paginated groups with
  selection, bulk delete and dismiss
- /delete: storage cards, delete by J-ID with preview grid, temp cleanup
- /users: staff directory with role selects and avatar J-ID inputs
- Nav + command palette entries; top-bar avatar; upload cards and the
  metadata modal show library visual matches
2026-09-17 12:49:10 -05:00
JakeBreath 6962e483fc Async Download to Library with progress; Download to client
Backend:
- DownloadTask model + background thread runner: streams the file with
  progress (%, bytes, speed) and a cancel flag, then indexes it, names it
  J-<id>.<ext> and applies the e621 metadata
- DownloadTaskViewSet (create/retrieve/cancel) replaces the synchronous
  endpoint; the status footer's worker counts now reflect download jobs
- Client download proxy (/api/online/file/) streams an e621 original to
  the browser with Content-Disposition: attachment, restricted to the
  configured e621 CDN hosts so it cannot be used as an open proxy

Frontend:
- Online detail: progress bar with percentage, transferred size, speed
  and cancel while downloading; success links to the new J-ID
- New 'Download to client' button available to everyone (guests too)
2026-09-17 12:06:15 -05:00
JakeBreath bf00cf36a2 Name uploaded and downloaded files J-<id>.<ext>
Files added through the upload pipeline and Download to Library are
renamed to their J-ID right after indexing, so every new library file is
traceable by its identifier (scanned files keep their existing names).
index_file now returns the created location so callers can rename it;
the location record is updated to the new path.
2026-09-17 11:56:01 -05:00
JakeBreath 1086beb974 Upload board: dismiss all, real previews for indexed records
- 'dismiss all' clears every indexed record at once
- Indexed cards show the actual file preview: completed records now get a
  signed library media URL (raw for images, thumbnail for videos) so
  <img>/<video> tags can load it, including items hidden from guests
- Media raw/thumbnail endpoints accept the signature for anonymous
  requests and fall back to the normal guest-filtered path otherwise
- Guest blacklist keeps a persistent Redis mirror: an expired TTL or an
  unreachable e621 keeps the last successful list instead of falling
  back to the small local list
2026-09-17 11:29:56 -05:00
JakeBreath b71ec729e0 Enrich IQDB candidates; link downloads the e621 original
- IQDB responses carry no preview/file data, so candidates only showed an
  ID; the SPA now enriches them with one batched posts lookup (preview,
  rating, score, favourites, dimensions, tag preview)
- Candidate tiles are selectable instead of instantly resolving: picking
  one shows its info and an explicit 'Link selected post' button
- Linking a post now fetches the e621 original into the library and
  drops the staged upload; when the staged file's MD5 already equals the
  post's file, the staged copy is moved instead (identical bytes)
- Keep the file URL in stored e621 metadata; sanitize the new candidate
  fields server-side
2026-09-17 11:24:01 -05:00
JakeBreath 7deb6084b6 Fix staged upload previews and allow WebP
- Staged files are now served through a signed URL (Django signing, 24h)
  so <img>/<video> tags can load previews without an Authorization
  header; the file endpoint accepts header auth or a valid signature,
  rejects tampered signatures, and still scopes access to the owner
- Serializer responses now carry the request context so URLs are signed
  per user
- Add .webp to the allowed extensions (backend + upload hint)
2026-09-17 11:17:04 -05:00
JakeBreath d0e2901c92 Upload pipeline: staging, MD5 auto-match, IQDB, three-column board
Backend:
- TempUpload model: staged files (pending / visual_match / completed /
  error) with resolution, e621 payload, custom metadata and IQDB data
- Files land in a temp folder and only move into the watched library
  folder once resolved; duplicates resolve immediately without a copy
- Endpoints: stage (multipart), list, retrieve, temp file, IQDB save,
  resolve (link to post or custom metadata), discard/dismiss
- cleanup_temp_uploads command for old staged files
- Replaces the old direct-to-library upload endpoint

Frontend:
- Upload page is now a three-column board (Pending & Unmatched /
  Visual Similarity Detected / Auto-uploaded & Indexed)
- After upload: MD5s are batch-checked against e621 and matches
  auto-complete with full post metadata; remaining files run through
  IQDB and move to the similarity column when candidates exist
- Metadata modal with IQDB candidates, post-ID linking and custom
  tags/rating/notes; discard and dismiss actions
- e621 client gains fetchPostsByMd5 and iqdbSearch helpers

Roadmap updated with the completed upload items.
2026-09-17 11:12:03 -05:00
JakeBreath ebac3ac922 Keep e621 metadata on downloaded posts; clear up account roles
Roles:
- JakeBreathild is now staff + superuser (the real account); the 'jake'
  smoke-test account was demoted to a regular user
- /me exposes is_superuser and the account page shows an admin badge

e621 metadata:
- MediaItem gains e621_post_id and e621_data (trimmed post payload:
  tags by category, rating, score, favourites, comments, sources,
  description, pools, relationships, file info, uploader)
- Download to Library accepts the post payload from the SPA and stores
  it; the item's custom rating is seeded from the e621 rating when empty
- Library detail shows an e621 metadata card: link to the in-app post,
  score/favourites/comments, taxonomy-coloured tags, DText description,
  sources and pools; grid cards get an e621 badge and fall back to the
  e621 rating for their colour (display_rating)
- Guest visibility now also considers e621 tags, so downloaded explicit
  content is hidden from anonymous visitors
2026-09-17 10:27:06 -05:00
JakeBreath e5cc63b0cc Phase 3: J-IDs, ownership, roles, guest safety, adaptive detail, download
Backend:
- User.role (user/uploader/staff) with can_upload; uploads and downloads
  gated to uploader+; owners and staff can edit their items
- MediaItem.uploaded_by plus J-<id> identity (serializer, admin,
  scan_files --user, first superuser as default owner)
- API resolves J-<id>, bare numeric ids and MD5s; neighbors and lookup
  return j_ids
- Guest safety: mirror e621's anonymous default blacklist into Redis
  (parses comments, negations and wildcards), flag hidden_from_guests
  and filter lists, details and lookups for anonymous users
- POST /api/online/downloads/ writes an e621 file into the watched
  folder and indexes it for the uploader
- MariaDB + Redis via docker compose (host ports 3307/6380), PyMySQL
  driver shim, Redis cache replacing the file cache; SQLite data
  dumped and loaded into MariaDB

Frontend:
- Single /detail/:itemId route with an adaptive shell: J-<id> renders
  the library item, bare numbers render the e621 post
- Legacy /view/<md5> and /online/view/<id> redirect to canonical URLs
- Cards expose J-IDs; library custom-data editor is read-only for
  non-owners
- Role gating: Upload hidden/blocked for regular users, account shows
  the role, guest hint on the library
2026-09-17 10:16:45 -05:00
JakeBreath 04448bf615 Online browser + post detail: SPA talks to e621 directly
Backend:
- POST /api/files/lookup/ reports which MD5s are already in the library

Frontend:
- e621 client extended with post/tag/favorite types and helpers: post
  search, post detail, batch posts by id, tag autocomplete, toggle
  favorite
- /online: tag search with autocomplete, post grid with rating colors
  and in-library badges, numbered pagination, page tag cloud, blacklist
  panel and filtered counts from /users/me.json, anonymous hint
- /online/view/🆔 media viewer (sample or original, video support),
  taxonomy-colored tags by category, specs sheet, favorite/unfavorite,
  description, sources, pools and parent/children thumbnails
- Shared CollapsibleSidebar extracted from the library page; Online
  added to the nav, command palette and sidebar toggle
2026-09-17 09:00:17 -05:00
JakeBreath a09ab8a160 Shell shortcuts: /, D, [ ], Ctrl+K palette + neighbors endpoint
Backend:
- GET /api/files/{md5}/neighbors/ returns previous/next items in the
  current ordering (name, size, created_at) for keyboard navigation

Frontend:
- / focuses the library search input
- D downloads the file on the detail page
- [ and ] navigate to the previous/next item; hint shown on the page
- Ctrl/Cmd+K opens a command palette (navigation, toggle filters,
  focus search, log in/out); Escape closes it
- Sort order now persists alongside the other library filters so
  prev/next stays consistent
2026-09-17 08:33:22 -05:00
JakeBreath 555c25d77b M0: scaffold SPA + API monorepo
Backend (Django 6.1 + DRF):
- Token auth with a custom User model (register/login/logout/me)
- Library models (MediaItem, MediaLocation) and REST endpoints
- File list/detail with search, rating filter, sorting, pagination
- Multipart upload with optional rating/tags/notes
- Range-aware media serving (video seeking) and ffmpeg thumbnails
- scan_files management command for the watched folder

Frontend (React 19 + Vite + TypeScript):
- Catppuccin Mocha design tokens from the design docs
- App shell, token persistence, protected routes
- Library grid with filters, file detail with custom data editor
- Upload page with per-file progress via XHR
- Dev proxy to the Django API
2026-09-17 07:56:47 -05:00