Commit Graph
13 Commits
Author SHA1 Message Date
JakeBreath 3a07481dfc Wait for all uploads, then batch MD5 -> visual -> IQDB with bulk links
Batching must not start while files are still being uploaded, and the MD5
phase must move a whole chunk at once instead of one resolve per file:

- the upload queue drains completely first (failed uploads included) before
  any matching starts;
- phase 1 asks e621 for every md5 (75 per posts.json request), builds the
  md5 -> post map from the response, and sends the matches to the new
  POST /api/uploads/link-bulk/ action, so a whole 75-file chunk moves into
  Indexed in a single board update;
- link-bulk indexes the staged file directly when the post's MD5 matches
  (identical bytes), so there is no per-file download round trip;
- phase 2 runs local visual similarity for whatever stayed pending, phase 3
  the IQDB queue.

Verified end to end with real e621 files: one md5 query for the batch, one
link-bulk call, both matching files flipped to Indexed together, then the
visual and IQDB phases. 23 library tests green (link-bulk, visual phase,
deferred visual matching).
2026-09-19 11:10:24 -05:00
JakeBreath e2bf1c457f Make upload processing phase-based: MD5 -> visual -> IQDB over the batch
Uploads were doing md5 + local visual matching inside the upload request
(backend create) while the frontend later ran its own e621 MD5 pass, so the
pipeline looked interleaved per file. Now every step is a phase applied to
the whole batch in order:

1. upload (fast: md5 + exact-duplicate check only),
2. e621 MD5 lookup, 75 md5: metatags per posts.json request,
3. local visual similarity, one file at a time via the new
   POST /api/uploads/<id>/visual-match/ action,
4. IQDB through the existing serial queue.

The board shows the active phase with its own progress bar (e621 MD5 in
peach, visual in lavender, IQDB in teal) and every step updates the staged
list as it lands. Verified from a headless run: one batched posts.json
request for 10 files, then 10 visual-match calls, then IQDB.
2026-09-19 10:25:57 -05:00
JakeBreath a761def65e Scrollable upload grid and bulk rating for the pending backlog
Upload board:
- the tile grid no longer re-sorts itself as files finish (that reshuffled
  the list under the cursor); it keeps insertion order, uses auto-fill tiles
  of ~150px so they hold a readable size, scrolls inside a 60vh area and no
  longer chains the page scroll (overscroll-contain);
- the files currently in flight are pinned in a small live strip above the
  grid (name, percent, bar) so progress stays visible while the grid is
  scrolled with hundreds of tiles.

Bulk rating: a 'bulk rate' button in the Pending & Unmatched header opens a
large modal with Safe/Questionable/Explicit pills, a tickable thumbnail grid
(Select all / Clear) and one confirm that moves every selected upload into
the library with that rating. Backed by POST /api/uploads/resolve-bulk/
(temp_ids + rating, own rows only): each staged file is resolved as a custom
entry (keeps its staged tags/notes), and already-completed or foreign ids are
reported per entry instead of failing the whole batch. Built for the
358-file backlog.

Tests: 4 bulk-resolve tests (resolution with the rating, input validation,
foreign ids untouched, mixed completed+pending) — full backend suite 53
green. Verified live end to end: staged a file, bulk-resolved it as 'q', saw
J-96 created with that rating, then removed the item, temp row and test
token.
2026-09-18 21:50:22 -05:00
JakeBreath 39307cb141 Unpaginate staged uploads and make big upload batches visible
The upload board partitions /api/uploads/ into Pending / Visual similarity /
Auto-uploaded, but the endpoint was paginated at 48 — a 69-file batch
silently lost 21 entries, and the similarity sweep (which reads the same
list back after uploading) only ever saw the first page. The staged-upload
list is now unpaginated: it is a transient per-user set, still limited to
the caller's rows and the uploader role. The page takes a plain array.

Watching progress with dozens of files was also poor:
- the queue uploads three files at a time instead of strictly one at a time;
- the Uploads section now shows a batch bar and 'n/m uploaded · x%' next to
  the count, so the overall progress never scrolls out of sight;
- entries are ordered active-first (uploading, queued, failed, done) so the
  file being uploaded is always at the top of the grid;
- tiles are larger (4 columns at lg instead of 5);
- the header reads 'Uploading n/m…' and 'Checking n file(s) against IQDB…'
  instead of a bare spinner.

Tests: staged-upload list unpaginated past 48, per-user, uploader-only
(3 new; full suite 49 green). Live-checked the bare-array response.
2026-09-18 19:58:17 -05:00
JakeBreath 99f617d296 Footer storage/backend display and a staff role that actually grants staff
Footer:
- Left is now 'Backend Storage:' with a capacity bar (blue, peach at 80%,
  red at 95% per DESIGN.md) and a used/total/free tooltip; the watched
  folder path is no longer printed. /api/status/ returns a compact storage
  summary instead of the path (the full storage page still shows paths to
  authenticated users).
- Centre shows the backend API origin (empty = same origin). Staff get a
  link to /setup to point the browser elsewhere; everyone else sees it as
  plain text. The Account 'Backend connection' card is gone — this is
  installation plumbing, not a per-user setting.
- Design spec updated to match.

Staff role:
- The custom role did nothing on several endpoints that only accepted
  Django's is_staff/is_superuser. One canonical check now exists:
  User.is_app_staff (superuser, Django staff, or the staff role), used by
  the stats/users APIs, item object permissions, can_delete, upload/
  similarity/download/match querysets, and the management commands
  (which also pick staff-role accounts for e621 sync/match and file
  ownership).

Verified with a role-only staff account (is_staff/is_superuser false):
stats/users 200, all 32 downloads + 2 scans visible, others' items
editable; the same account as role=user gets 403 for all of those.
2026-09-17 23:24:24 -05:00
JakeBreath 16907c39ca Support cross-origin frontends alongside same-origin setups
- django-cors-headers with env-driven CORS_ALLOWED_ORIGINS,
  CORS_ALLOW_ALL_ORIGINS, CORS_ALLOW_CREDENTIALS and CSRF_TRUSTED_ORIGINS;
  same-origin traffic is unaffected and a disallowed origin gets no CORS
  headers. Token auth needs no cookies, so credentials stay off by default.
- TRUST_PROXY_HEADERS=true lets a TLS-terminating proxy supply
  X-Forwarded-Proto/Host for correct absolute URLs.
- API media URLs (raw/thumbnail/upload/similarity/staged previews) are now
  absolute, built from the request host, so <img>/<video>/fetch() keep
  working when the SPA is served from another origin. Signed URLs are still
  per-user; nothing is stored in the DB.
- The SPA gains VITE_API_BASE (build-time, empty = same-origin) applied by
  a small apiUrl() helper used for XHR/fetch and the few URL fallbacks.

Verified with a throwaway instance: preflight and GET responses carry the
allowed origin, foreign origins get nothing, media GETs include CORS for
cross-origin fetch(), and payload URLs use the request host (dev :8000
unchanged).
2026-09-17 22:50:12 -05:00
JakeBreath 09405d1a0f Match local files to e621: MD5 lookups, manual links, batch scans
- MediaItem gains e621_match_status (unknown/matched/not_found/deleted)
  and e621_checked_at, backfilled for existing matched items.
- Server-side e621 client (apps/library/e621.py) using the user's stored
  credentials, throttled to 2 req/s, with typed errors.
- Matching service: MD5 lookup, manual post linking (flags MD5
  mismatches), unlink, metadata refresh, deleted-post detection.
- Detail actions POST /api/files/J-x/match/ and /unlink/ (uploader or
  staff only).
- Background library scans: MatchTask + /api/matches/ with missing/all
  scopes, progress polling, cancel and stale-task reaping; the scan
  counts toward the footer's Active Workers. Same pass available as
  manage.py match_e621 for cron.
- Library gains not_found/deleted status filters; the detail page adds
  an e621 match card (check / link by post ID / unlink) and the metadata
  card warns when a post was deleted on e621.
2026-09-17 13:41:08 -05:00
JakeBreath db74f7ab18 Fix hidden library items not rendering in the browser
Items flagged hidden_from_guests (blacklisted tags) returned 404 for
<img> requests since tags cannot send the auth header. The API now
exposes signed raw_url/thumbnail_url fields (mirroring upload previews
and avatars), and the SPA uses them in the gallery, detail view,
duplicates and delete screens, and upload visual matches.
2026-09-17 13:25:20 -05:00
JakeBreath cd490b0a23 Duplicates, delete & storage, users page with J-ID avatars
Backend:
- Perceptual hashes (aHash/dHash/pHash/wHash via imagehash, no imgdd)
  stored on items, computed on upload/download and by the new
  compute_visual_hashes command
- Duplicates API: exact duplicates (multi-location items), visual matches
  for one item, union-find similarity groups with pagination
- Delete API with ownership/staff checks, per-item and per-copy deletion,
  watched-folder path validation; storage overview and temp cleanup;
  file list accepts j_ids batches
- Staged uploads are flagged visual_match with their library matches
  (threshold via VISUAL_MATCH_THRESHOLD)
- Staff users API: list with upload counts, set role and avatar by J-ID;
  User.avatar FK with signed avatar URLs
- Download threads close their DB connection and stale tasks are reaped,
  keeping behaviour Gunicorn-friendly

Frontend:
- /duplicates: exact duplicate groups with per-copy delete, visual
  similarity controls, search similar to a J-ID, paginated groups with
  selection, bulk delete and dismiss
- /delete: storage cards, delete by J-ID with preview grid, temp cleanup
- /users: staff directory with role selects and avatar J-ID inputs
- Nav + command palette entries; top-bar avatar; upload cards and the
  metadata modal show library visual matches
2026-09-17 12:49:10 -05:00
JakeBreath bf00cf36a2 Name uploaded and downloaded files J-<id>.<ext>
Files added through the upload pipeline and Download to Library are
renamed to their J-ID right after indexing, so every new library file is
traceable by its identifier (scanned files keep their existing names).
index_file now returns the created location so callers can rename it;
the location record is updated to the new path.
2026-09-17 11:56:01 -05:00
JakeBreath b71ec729e0 Enrich IQDB candidates; link downloads the e621 original
- IQDB responses carry no preview/file data, so candidates only showed an
  ID; the SPA now enriches them with one batched posts lookup (preview,
  rating, score, favourites, dimensions, tag preview)
- Candidate tiles are selectable instead of instantly resolving: picking
  one shows its info and an explicit 'Link selected post' button
- Linking a post now fetches the e621 original into the library and
  drops the staged upload; when the staged file's MD5 already equals the
  post's file, the staged copy is moved instead (identical bytes)
- Keep the file URL in stored e621 metadata; sanitize the new candidate
  fields server-side
2026-09-17 11:24:01 -05:00
JakeBreath 7deb6084b6 Fix staged upload previews and allow WebP
- Staged files are now served through a signed URL (Django signing, 24h)
  so <img>/<video> tags can load previews without an Authorization
  header; the file endpoint accepts header auth or a valid signature,
  rejects tampered signatures, and still scopes access to the owner
- Serializer responses now carry the request context so URLs are signed
  per user
- Add .webp to the allowed extensions (backend + upload hint)
2026-09-17 11:17:04 -05:00
JakeBreath d0e2901c92 Upload pipeline: staging, MD5 auto-match, IQDB, three-column board
Backend:
- TempUpload model: staged files (pending / visual_match / completed /
  error) with resolution, e621 payload, custom metadata and IQDB data
- Files land in a temp folder and only move into the watched library
  folder once resolved; duplicates resolve immediately without a copy
- Endpoints: stage (multipart), list, retrieve, temp file, IQDB save,
  resolve (link to post or custom metadata), discard/dismiss
- cleanup_temp_uploads command for old staged files
- Replaces the old direct-to-library upload endpoint

Frontend:
- Upload page is now a three-column board (Pending & Unmatched /
  Visual Similarity Detected / Auto-uploaded & Indexed)
- After upload: MD5s are batch-checked against e621 and matches
  auto-complete with full post metadata; remaining files run through
  IQDB and move to the similarity column when candidates exist
- Metadata modal with IQDB candidates, post-ID linking and custom
  tags/rating/notes; discard and dismiss actions
- e621 client gains fetchPostsByMd5 and iqdbSearch helpers

Roadmap updated with the completed upload items.
2026-09-17 11:12:03 -05:00