Author SHA1 Message Date
JakeBreath 7ca84f4fea Point desktop updates at the Gitea release feed
CI / Backend tests (push) Successful in 2m41s
CI / Frontend build & lint (push) Successful in 25s
The updater now resolves the newest non-draft desktop-v* release through the
Gitea API at check time (J621_UPDATE_REPO, lowercase because the API path is
case-sensitive), picks the platform's latest*.yml asset and uses that release
as a generic electron-updater feed; J621_UPDATE_URL still overrides
everything. Verified against the live API: release picked, yml fetched,
artifact HEAD 200.

electron-builder's publish.url is now metadata only (still needed so the
build emits latest*.yml). Docs updated: CD release assets are the feed, the
website /desktop/ feed only matters for installs before 0.1.2.
2026-09-23 22:00:49 -05:00
JakeBreath 5f237aa2e3 Widen the header and collapse blacklist listings by default
- the header row is no longer capped at 1600px, so the nav hugs the left
  edge and the status/account controls the right; main content and footer
  keep their width
- OnlinePage's blacklist chips and the Followed page's blacklisted-tag
  cloud start collapsed behind a count toggle (N tags / N entries)
2026-09-23 21:57:48 -05:00
JakeBreath 90ba2ecff7 Bump the desktop app to 0.1.2 2026-09-23 21:40:43 -05:00
JakeBreath 73bf4f9e38 Rework the metadata modal: fullscreen, inline sections, bigger previews
- near-fullscreen panel (up to 1400px / 92vh) with two independently
  scrolling columns; the left preview uses self-start so its border hugs the
  image instead of stretching to the modal height
- J-ID matches render as a larger tile grid (was 48px rows) and IQDB
  candidates get bigger tiles too
- Link to e621 post and Custom metadata are shown inline instead of behind
  tabs, each with its own heading
- new ThumbImage component: spinner while loading and a broken-image icon on
  error, with alt text removed so a pending tile never reads as 'J-7786'
2026-09-23 21:39:48 -05:00
JakeBreath 7c2569522f Make thumbnail generation atomic and warm it on import
- write thumbnails to a .part file and os.replace() them, so concurrent
  requests never read a half-written JPEG
- a stale thumbnail plus a vanished source no longer raises through the
  request (getmtime on a missing file returned 500); it falls back cleanly
- ensure_thumbnail(item) warms the preview when a file is indexed, keeping
  image decoding out of the request path
2026-09-23 21:35:09 -05:00
JakeBreath a62195ffce Re-sign visual-match thumbnails on every detail fetch
Match rows stored a signed URL minted when the scan ran, so it aged out (or
used the pre-stable signing scheme) and the modal showed broken tiles even
after legacy signatures were fixed. Rows now carry item_id/j_id and the
detail serializer mints a fresh thumbnail URL per request; matches whose
item no longer exists are dropped.
2026-09-23 21:33:20 -05:00
JakeBreath c73a81a5f4 Fix 500 on legacy signed URLs
A TimestampSigner value is an HMAC over 'payload:timestamp', so a plain
Signer's HMAC check accepts it and the embedded timestamp then reached the
JSON decoder, raising JSONDecodeError (not BadSignature) and surfacing as a
500. That broke every stored visual-match thumbnail URL minted before the
stable scheme, so the J-ID match tiles never loaded on prod.

Detect the legacy shape by its extra separator and verify it with
TimestampSigner; malformed input returns None instead of raising.
2026-09-23 21:31:34 -05:00
JakeBreath af378e7d71 Update the roadmap for cacheable media and session-only completions
CI / Backend tests (push) Successful in 2m29s
CI / Frontend build & lint (push) Successful in 25s
2026-09-23 18:29:49 -05:00
JakeBreath 59f397a96c Show indexed uploads as session notifications, add a Failed column
The board rendered every persisted completed row with a dismiss button and a
dismiss-all that sent the whole column to the 1000-id bulk endpoint. Now that
the backend deletes completed rows and reports them through the status feed:

- the Auto-uploaded & Indexed column renders the live feed only; a reload or
  leaving the page forgets them, and there is nothing to dismiss (just a
  client-side clear)
- duplicates complete during staging and get their card immediately from the
  upload response
- failures get their own persisted column with per-card retry and discard
- bulk discard chunks requests at 500 ids, so backlogs over the server's
  1000-id cap are still removable in one action
2026-09-23 18:28:35 -05:00
JakeBreath 9ababb8b48 Stop storing completed uploads; announce them through a live feed
Every auto-matched, duplicate or manually resolved upload left a completed
TempUpload row on the board until it was dismissed by hand, so the rows
accumulated without bound and the bulk dismiss (capped at 1000 ids) failed
once there were more. The original app never stored these: they are
notifications, not records.

- complete_temp_upload now appends {filename, J-ID, resolution, post} to a
  bounded recent_completions feed on UploadRun and deletes the staged row
- staging duplicates never create a board record either; the create response
  carries the J-ID and preview so the SPA can show the card immediately
- status_payload returns the feed (newest first, signed thumbnails) for the
  live board; finalize_round counts deleted matches in processed
- resolve/link-bulk return synthetic completion payloads
- migration 0012 adds the field and purges the existing completed backlog
  (and any stray staged files) on deploy
2026-09-23 18:24:38 -05:00
JakeBreath 37085c5dac Make media URLs stable and cacheable, add real image thumbnails
Signed media URLs embedded the current second (TimestampSigner), so every
API response re-minted every raw/thumbnail/staged URL and the browser
re-downloaded each file on every poll or navigation. Responses also carried
no cache headers at all.

- sign with a plain Signer plus a bucket-quantized exp (7d TTL, 24h bucket),
  so a URL is byte-identical across responses and rotates once a day; legacy
  TimestampSigner URLs stay accepted for one release
- add a v=<md5> version parameter to library media URLs so replacing a file
  under the same J-ID (the optimize flow) busts caches exactly when needed
- serve_file now sends ETag/Last-Modified and a private Cache-Control and
  answers conditional requests with 304; library media gets max-age 6d +
  immutable, staged/similarity files 1h
- build cached 480px JPEG thumbnails for images (Pillow, keyed by MD5 under
  MEDIA_ROOT/thumbs) instead of serving full-size originals through the
  thumbnail endpoint; the library grid uses thumbnail_url for images too
2026-09-23 18:13:52 -05:00
JakeBreath ecac4cb8b4 Fix the release asset upload for names with spaces
CI / Backend tests (push) Successful in 2m4s
CI / Frontend build & lint (push) Successful in 23s
curl exit 3 (malformed URL) on 'J621 Setup 0.1.1.exe': percent-encode the
asset name in the query string.

Also replace assets instead of skipping them on re-runs: NSIS builds are
not bit-reproducible, so latest.yml/latest-linux.yml must reference the
installers produced by the same run. Existing assets are deleted by id
before the fresh upload.
2026-09-23 07:14:19 -05:00
JakeBreath 2eb7af0d41 Fix the CD wine build and add per-part toggles
CI / Backend tests (push) Successful in 2m13s
CI / Frontend build & lint (push) Successful in 20s
The Windows NSIS step failed under wine for two reasons: no X display
(nodrv_CreateWindow) and missing 32-bit libraries (failed to load
syswow64\ntdll.dll). The job now installs xvfb + wine32:i386 and runs the
desktop build under xvfb-run, with Gecko/Mono lookups disabled.

Also add `images` and `desktop` dispatch inputs so either half of the CD
can be skipped (e.g. desktop-only or images-only releases).
2026-09-23 00:06:55 -05:00
JakeBreath 7cecfeabc6 Use a minimal registry PAT and the job token for releases
CI / Backend tests (push) Successful in 2m23s
CI / Frontend build & lint (push) Successful in 22s
Gitea's container registry rejects the automatic job token
(go-gitea/gitea#23642 is still open), so the image push keeps a PAT with
only the write:package scope; a preflight step fails clearly when the
REGISTRY_USER/REGISTRY_TOKEN secrets are missing. Release creation needs
no PAT: the desktop job asks for contents: write on the job token.
2026-09-22 23:35:55 -05:00
JakeBreath ed6178d12e Make Actions token permissions explicit
The automatic job token creates the desktop release, so the CD desktop job
asks for contents: write; the images job keeps contents: read and asks for
packages: write so the job token can stand in for the scoped registry PAT.
CI stays read-only. The registry token itself remains a write:package-only
PAT (verified login + pull).
2026-09-22 23:30:37 -05:00
JakeBreath 8f9656ac0e Add the manual CD release workflow
One dispatch builds and pushes both images and builds the desktop packages
into a Gitea release (desktop-v<version>, installers + latest*.yml attached,
idempotent on re-run). The live update feed stays a deploy-host operation:
CI has no SSH key for jakerasp, so push_desktop.sh --no-build remains the
way to publish it.

Repo secrets REGISTRY_USER/REGISTRY_TOKEN are set, so the image push uses
the Gitea registry credentials directly.
2026-09-22 23:23:05 -05:00
JakeBreath 72fc42217f Fix CI for the user-scoped runners
- ci.yml: connect to the test MariaDB as root so Django creates the test
  database itself (no client install/grant step), and drop actions/cache
  (cache: pip/npm): Gitea's cache service hangs the job on restore/save.
- publish.yml: prefer the REGISTRY_USER/REGISTRY_TOKEN secrets (as on other
  repos) and fall back to the automatic Actions token.
- AGENTS.md: note the CI layout, the runner labels and the cache caveat.
2026-09-22 23:12:56 -05:00
JakeBreath d9c1e9e521 Add CI and manual image publishing workflows
CI / Backend tests (push) Successful in 12m54s
CI / Frontend build & lint (push) Failing after 4m59s
- .gitea/workflows/ci.yml: on every push/PR, run Django checks + the full
  backend suite against MariaDB/Redis services and the frontend
  lint/type-check/build. Runs on the nitro-ci runner (ubuntu-latest).
- .gitea/workflows/publish.yml: manual dispatch; multi-arch build+push of
  both images as :latest and :<short-sha> with GIT_HASH baked in.
- push_*.sh: non-interactive registry login for CI (REGISTRY_USER/
  REGISTRY_TOKEN) and a PLATFORMS override.
2026-09-22 22:32:37 -05:00
JakeBreath e2697c0a78 Stop throttling signed media and ease the browser's e621 queue
Signed media URLs are fetched by <img>/<video> tags without an
Authorization header, so they were charged to the anonymous 120/min
bucket: past that, galleries and the fish-greeting download got 429 JSON
instead of image bytes. The raw/thumbnail/staged-file/similarity-file
actions are now exempt, and THROTTLE_ENABLED=false removes the general
anon+user limits for private/tailnet deployments (login/register/proxy
guards stay).

The SPA's e621 client also stops self-throttling so hard: 1s gap between
browsing calls (2.5s for the stricter IQDB endpoint) and a 15s cooldown
instead of 60s when e621 answers 429.
2026-09-22 22:32:37 -05:00
JakeBreath 474403ffe2 Upload updates 2026-09-21 09:01:01 -05:00
JakeBreath 98025e9e6d Prune old installers from the remote feed on push
rsync without --delete left every previous version on the server (the
screenshot showed 0.1.0 and 0.1.1 side by side). The push now removes
non-current installers over ssh first, so the remote feed mirrors the local
one whether the transfer uses rsync or tar.
2026-09-20 21:39:43 -05:00
JakeBreath 3183a3bece Prune old desktop builds from release/ on every build
build_desktop.sh now reads the version first and removes anything in
desktop/release/ that is not that version (plus the regenerated unpacked
trees), so a version bump never leaves old installers lying around — the
same rule push_desktop.sh applies to the feed.
2026-09-20 21:16:56 -05:00
42 changed files with 4214 additions and 1034 deletions
+197
View File
@@ -0,0 +1,197 @@
# J621 CD — manual release workflow (Actions tab -> "Run workflow").
#
# One dispatch does everything; each half can be skipped with the `images`
# and `desktop` inputs:
# * builds and pushes the backend + frontend images (multi-arch, :latest
# and :<short-sha>, GIT_HASH baked in for the version pill),
# * builds the desktop packages and attaches them (plus the update
# metadata) to the Gitea release tagged `desktop-v<package.json version>`.
#
# The desktop build also attaches the update metadata (latest*.yml) to the
# release; that is the desktop updater's feed, resolved through the Gitea API
# at check time (see desktop/README.md). The older website feed
# (deploy/data/desktop, served at /desktop/) is runtime state on the deploy
# host and only needed for installs before 0.1.2; it is refreshed with
# `deploy/push_desktop.sh` from a machine that can reach the deploy host.
#
# Registry login uses a repo PAT with the minimal write:package scope (the
# Gitea registry rejects the automatic job token, go-gitea/gitea#23642);
# release creation uses the automatic job token. Jobs run on the user-scoped
# nitro-ci runner (ubuntu-latest).
name: CD
on:
workflow_dispatch:
inputs:
images:
description: Build and push the Docker images
required: false
default: "true"
desktop:
description: Build the desktop release
required: false
default: "true"
platforms:
description: Image platforms (comma separated)
required: false
default: linux/amd64,linux/arm64
windows:
description: Also cross-build the Windows installer (needs wine, slow)
required: false
default: "false"
concurrency:
group: cd
cancel-in-progress: false
jobs:
images:
name: Build & push images
if: ${{ inputs.images != 'false' }}
runs-on: ubuntu-latest
# The Gitea container registry does not accept the automatic job token
# (go-gitea/gitea#23642 is still open), so the push uses a repo PAT with
# the minimal write:package scope. Releases use the job token instead.
permissions:
contents: read
env:
REGISTRY_USER: ${{ secrets.REGISTRY_USER }}
REGISTRY_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
PLATFORMS: ${{ inputs.platforms || 'linux/amd64,linux/arm64' }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Check the registry credentials
run: |
if [ -z "$REGISTRY_USER" ] || [ -z "$REGISTRY_TOKEN" ]; then
echo "Set the REGISTRY_USER and REGISTRY_TOKEN repo secrets" >&2
echo "(a PAT with the write:package scope)." >&2
exit 1
fi
- name: Register binfmt (multi-arch builds)
run: docker run --privileged --rm tonistiigi/binfmt --install all
- name: Build & push both images
run: |
set -euo pipefail
SHA="$(git rev-parse --short HEAD)"
echo "Publishing $SHA for $PLATFORMS"
PLATFORMS="$PLATFORMS" ./deploy/push_frontend.sh "$SHA"
PLATFORMS="$PLATFORMS" ./deploy/push_backend.sh "$SHA"
desktop:
name: Desktop release
if: ${{ inputs.desktop != 'false' }}
runs-on: ubuntu-latest
# Creating the release and uploading its assets uses the automatic job
# token, so it needs write access to the repository's releases.
permissions:
contents: write
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: "22"
- name: Install packaging tools
run: |
sudo apt-get update
sudo apt-get install -y --no-install-recommends fakeroot libarchive-tools
if [ "${{ inputs.windows }}" = "true" ]; then
# electron-builder runs the 32-bit NSIS installer under wine to
# build the uninstaller: that needs a virtual display (Xvfb) and
# 32-bit wine libraries.
sudo dpkg --add-architecture i386
sudo apt-get update
sudo apt-get install -y --no-install-recommends xvfb wine wine32:i386
fi
- name: Install frontend + desktop dependencies
run: |
npm --prefix frontend ci --no-audit --no-fund
npm --prefix desktop ci --no-audit --no-fund
- name: Build desktop packages
env:
# Keep wine from trying to fetch Gecko/Mono on first run.
WINEDLLOVERRIDES: mscoree,mshtml=
run: |
if [ "${{ inputs.windows }}" = "true" ]; then
xvfb-run -a ./deploy/build_desktop.sh --all
else
./deploy/build_desktop.sh --linux
fi
- name: Add the Gitea release
env:
GITEA_TOKEN: ${{ secrets.GITEA_TOKEN || secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
VERSION="$(node -p "require('./desktop/package.json').version")"
TAG="desktop-v$VERSION"
API="${{ github.server_url }}/api/v1/repos/${{ github.repository }}"
AUTH="Authorization: token $GITEA_TOKEN"
NOTES="$(printf 'J621 desktop %s\n\n' "$VERSION"
cd desktop/release
sha256sum ./*.deb ./*.pkg.tar.zst ./*.exe 2>/dev/null || true)"
RELEASE_ID="$(curl -sf -H "$AUTH" "$API/releases/tags/$TAG" \
| python3 -c 'import json,sys; print(json.load(sys.stdin).get("id",""))' \
2>/dev/null || true)"
if [ -z "$RELEASE_ID" ]; then
echo "Creating release $TAG"
PAYLOAD="$(python3 - "$TAG" "${{ github.sha }}" "$NOTES" <<'PY'
import json, sys
print(json.dumps({
"tag_name": sys.argv[1],
"name": sys.argv[1],
"body": sys.argv[3],
"target_commitish": sys.argv[2],
}))
PY
)"
RELEASE_ID="$(curl -sf -X POST -H "$AUTH" \
-H "Content-Type: application/json" -d "$PAYLOAD" "$API/releases" \
| python3 -c 'import json,sys; print(json.load(sys.stdin)["id"])')"
else
echo "Release $TAG already exists (id $RELEASE_ID); attaching missing files."
fi
EXISTING="$(curl -sf -H "$AUTH" "$API/releases/$RELEASE_ID/assets" \
|| echo '[]')"
for FILE in desktop/release/*"$VERSION"*.deb \
desktop/release/*"$VERSION"*.pkg.tar.zst \
desktop/release/latest-linux.yml \
desktop/release/latest.yml \
desktop/release/*"$VERSION"*.exe \
desktop/release/*"$VERSION"*.exe.blockmap; do
[ -e "$FILE" ] || continue
NAME="$(basename "$FILE")"
# Replace the asset when it is already there: latest*.yml must
# reference the installers built by *this* run (NSIS builds are
# not bit-reproducible), so old copies are deleted first.
ASSET_ID="$(printf '%s' "$EXISTING" | python3 -c '
import json, sys
name = sys.argv[1]
print(next((str(a["id"]) for a in json.load(sys.stdin) if a["name"] == name), ""))
' "$NAME")"
if [ -n "$ASSET_ID" ]; then
echo " replacing $NAME"
curl -sf -X DELETE -H "$AUTH" \
"$API/releases/$RELEASE_ID/assets/$ASSET_ID" >/dev/null
else
echo " attaching $NAME"
fi
# Names like "J621 Setup 0.1.1.exe" contain spaces: encode them
# or curl refuses the URL (exit 3).
ENCODED="$(python3 -c 'import sys, urllib.parse; print(urllib.parse.quote(sys.argv[1]))' "$NAME")"
curl -sf -X POST -H "$AUTH" -H "Content-Type: application/octet-stream" \
--data-binary @"$FILE" "$API/releases/$RELEASE_ID/assets?name=$ENCODED" >/dev/null
done
echo "Release: ${{ github.server_url }}/${{ github.repository }}/releases/tag/$TAG"
+101
View File
@@ -0,0 +1,101 @@
# J621 CI — runs on every push (and pull request): Django checks + the full
# backend test suite against MariaDB/Redis, and the frontend type-check,
# lint and production build.
#
# Runner: the "nitro-ci" act_runner with the custom `ubuntu-latest` label.
name: CI
# Tests and builds only need to read the repository; the automatic job token
# stays read-only.
permissions:
contents: read
on:
push:
branches: ["**"]
tags-ignore: ["**"]
pull_request:
workflow_dispatch:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
backend:
name: Backend tests
runs-on: ubuntu-latest
services:
mariadb:
image: mariadb:11.4
env:
MARIADB_ROOT_PASSWORD: root
MARIADB_DATABASE: j621
MARIADB_USER: j621
MARIADB_PASSWORD: j621
options: >-
--health-cmd="healthcheck.sh --connect --innodb_initialized"
--health-interval=5s
--health-timeout=5s
--health-retries=12
redis:
image: redis:7-alpine
options: >-
--health-cmd="redis-cli ping"
--health-interval=5s
--health-timeout=5s
--health-retries=12
env:
# Connect as root so Django can create the test database itself;
# everything else mirrors the development defaults.
DB_HOST: mariadb
DB_PORT: "3306"
DB_NAME: j621
DB_USER: root
DB_PASSWORD: root
REDIS_URL: redis://redis:6379/1
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.14"
- name: Install backend dependencies
run: pip install -r backend/requirements.txt
- name: Django system checks
working-directory: backend
run: python manage.py check
- name: Backend tests
working-directory: backend
run: >-
python manage.py test
apps.core.tests
apps.library.tests
apps.follows.tests
apps.accounts.tests
frontend:
name: Frontend build & lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: "22"
- name: Install frontend dependencies
working-directory: frontend
run: npm ci
- name: Lint
working-directory: frontend
run: npm run lint
- name: Type-check & build
working-directory: frontend
run: npm run build
+17 -1
View File
@@ -49,6 +49,17 @@ Project constraints (do not regress):
- Periodic commands (follow syncs, similarity cleanup, guest blacklist
refresh) run in the composes' `scheduler` service — the backend image with
the j621-scheduler entrypoint, intervals via J621_*_EVERY. No host cron.
- CI/CD lives in .gitea/workflows: ci.yml runs on every push/PR (Django checks
+ the full backend suite against MariaDB/Redis service containers, frontend
lint/type-check/build); cd.yml is manual and builds/pushes both images
multi-arch plus the desktop packages (attached to the Gitea release
`desktop-v<version>`). Jobs run on the user-scoped runners: `ubuntu-latest`
on nitro-ci, `desktop` on msi-mortar-ci. Do not add actions/cache
(`cache: pip`/`npm`) to these workflows: Gitea's cache service hangs the job
on restore/save. Desktop updates read the release assets (latest*.yml) from
the Gitea API at check time; the older website feed (deploy/data/desktop)
only matters for installs before 0.1.2 and is refreshed with
`deploy/push_desktop.sh` from a machine with SSH to the deploy host.
- Security/permission tests live in backend/apps/core/tests and need a
one-time grant: GRANT ALL ON `test_j621`.* TO 'j621'@'%';
@@ -85,7 +96,12 @@ Security hardening (do not weaken):
SECRET_KEY (apps/accounts/crypto.py); rotating SECRET_KEY invalidates them
(and all signed media URLs), so users must re-enter the key.
- API throttles live in REST_FRAMEWORK (env-overridable): anon 120/min,
user 600/min, login 5/min, register 20/hour, e621_proxy 60/hour.
user 600/min, login 5/min, register 20/hour, e621_proxy 60/hour. Signed
media URLs (raw/thumbnail/staged-file/similarity-file actions) are exempt
on purpose: <img>/<video> tags fetch them without an Authorization header,
so a gallery would otherwise drain the anonymous bucket and get 429 JSON
instead of images. THROTTLE_ENABLED=false removes the anon+user limits for
private/tailnet deployments (the login/register/proxy guards stay).
- Only admins (superusers) may grant/revoke the staff role or delete
staff/admin accounts; staff manage regular/uploader accounts only.
- Storage, duplicates, delete, temp-clear, uploads and downloads require
+24 -10
View File
@@ -20,6 +20,13 @@ they land.
- [x] Tag cloud in the sidebar (click to search, hidden from guests for
blacklisted items)
- [x] Status filter (matched / not_found / deleted / custom / unknown)
- [x] **Cacheable media serving**
- [x] Stable signed URLs (7 d TTL, 24 h rotation) plus a `v=<md5>` cache
buster, so replacing a file under the same J-ID invalidates it
- [x] ETag/Last-Modified and a private `Cache-Control` (immutable for
versioned media), conditional 304s
- [x] Real 480 px image thumbnails (cached by MD5) instead of serving
full-size originals through the thumbnail endpoint
- [x] **Random image** endpoint and page: `GET /api/random/` (aliases
`/random`, `/random/`) with `rating=s,q,e` filters; fastfetch mode
@@ -95,18 +102,25 @@ Files now stage first and are resolved before entering the library.
- [x] `cleanup_temp_uploads` command for old staged files
- [x] **Auto-upload / auto-match**
- [x] MD5 computed on staging; exact duplicates resolve immediately
- [x] MD5 batch-checked against e621; matches auto-complete with post
metadata stored and rating seeded
- [x] **IQDB similarity on upload** (SPA-driven)
- [x] Automatic + manual IQDB checks with candidate posts
- [x] "Visual Similarity Detected" state with candidate picker
- [x] Perceptual-hash comparison against the library (staged uploads are
flagged with their library matches as soon as they land)
- [x] MD5 batch-checked against e621 in chunks of 75; matches auto-complete
with post metadata stored and rating seeded
- [x] **Background pipeline** (server-side)
- [x] Daemon-thread worker runs MD5 → visual similarity → IQDB for every
staged upload, so the work continues after the page or tab is closed
- [x] Durable progress (`UploadRun` + per-file phase flags) polled by the
shell indicator; rate-limited runs retry with backoff
- [x] Perceptual-hash comparison against the library, loaded once per batch
- [x] IQDB candidates stored with one batched enrichment request
- [x] Indexed uploads are announced through a bounded per-run feed and the
staged row is deleted; nothing persists on the board to dismiss
- [x] **Upload UI**
- [x] Three-column board: Pending & Unmatched / Visual Similarity Detected /
Auto-uploaded & Indexed
- [x] Four-column board: Pending & Unmatched / Visual Similarity Detected /
Auto-uploaded & Indexed (session feed) / Failed, private per user
(staff included)
- [x] Metadata modal (link to e621 post, IQDB candidates, custom metadata)
- [x] Per-file progress plus batch processing indicator
- [x] Per-file progress plus background pipeline status
- [x] Bulk actions: bulk rate and discard all (chunked past the API's
1000-id cap)
## 4. Staff tools
+70 -20
View File
@@ -36,6 +36,7 @@ from apps.library.models import (
TempUpload,
)
from apps.library.services import MEDIA_FILE_SALT
from apps.library.signing_urls import sign_payload
User = get_user_model()
@@ -122,21 +123,15 @@ class SecurityTestCase(TestCase):
return item
def old_signature(self, item, action="raw", age=3 * 86400):
"""A valid signature minted `age` seconds ago."""
real_time = signing.time
class Backdated:
def time(self):
return real_time.time() - age
try:
signing.time = Backdated()
return signing.dumps(
{"item": item.id, "user": self.users["sec-uploader"].id, "action": action},
salt=MEDIA_FILE_SALT,
"""A signed media URL whose expiry is `age` seconds in the past."""
return signing.Signer(salt=MEDIA_FILE_SALT).sign_object(
{
"item": item.id,
"user": self.users["sec-uploader"].id,
"action": action,
"exp": int(time.time()) - age,
}
)
finally:
signing.time = real_time
class GuestVisibilityTests(SecurityTestCase):
@@ -183,9 +178,9 @@ class GuestVisibilityTests(SecurityTestCase):
def test_authenticated_users_and_signed_urls_see_protected_items(self):
uploader = self.client_for("sec-uploader")
self.assertEqual(uploader.get(f"/api/files/J-{self.hidden.id}/").status_code, 200)
signed = signing.dumps(
signed = sign_payload(
{"item": self.hidden.id, "user": self.users["sec-uploader"].id, "action": "raw"},
salt=MEDIA_FILE_SALT,
MEDIA_FILE_SALT,
)
self.assertEqual(
self.guest.get(f"/api/files/J-{self.hidden.id}/raw/?sig={signed}").status_code,
@@ -193,9 +188,9 @@ class GuestVisibilityTests(SecurityTestCase):
)
def test_signature_integrity(self):
signed = signing.dumps(
signed = sign_payload(
{"item": self.hidden.id, "user": self.users["sec-uploader"].id, "action": "raw"},
salt=MEDIA_FILE_SALT,
MEDIA_FILE_SALT,
)
raw = f"/api/files/J-{self.hidden.id}/raw/"
thumbnail = f"/api/files/J-{self.hidden.id}/thumbnail/"
@@ -203,10 +198,29 @@ class GuestVisibilityTests(SecurityTestCase):
self.assertEqual(self.guest.get(f"{raw}?sig={signed[:-4]}AAAA").status_code, 404)
# Valid signature, wrong action.
self.assertEqual(self.guest.get(f"{thumbnail}?sig={signed}").status_code, 404)
# Expired signature (minted three days ago).
# Expired signature (expiry three days ago).
expired = self.old_signature(self.hidden)
self.assertEqual(self.guest.get(f"{raw}?sig={expired}").status_code, 404)
def test_signed_media_urls_are_stable_and_versioned(self):
"""The same item must keep the same URL across responses.
A per-second signature made browsers re-download every image on every
poll; the MD5 version parameter busts caches only when the file itself
changes (the optimize flow rewrites files under the same J-ID).
"""
item = self.visible
first = services.signed_media_url(item, self.users["sec-uploader"])
time.sleep(1.1)
second = services.signed_media_url(item, self.users["sec-uploader"])
self.assertEqual(first, second)
self.assertIn(f"v={item.md5}", first)
MediaItem.objects.filter(pk=item.pk).update(md5="b" * 32)
item.refresh_from_db()
self.assertNotEqual(
services.signed_media_url(item, self.users["sec-uploader"]), first
)
class RoleBoundaryTests(SecurityTestCase):
def test_non_uploader_is_read_only(self):
@@ -345,7 +359,13 @@ class PrivacyTests(SecurityTestCase):
)
self.assertEqual(self.guest.get(f"/api/uploads/{temp.id}/file/").status_code, 401)
self.assertNotIn(
str(temp.id), self.client_for("sec-uploader").get("/api/uploads/").content.decode()
str(temp.id),
self.client_for("sec-uploader").get("/api/uploads/").content.decode(),
)
# The board is per-user: staff only see their own staged uploads.
self.assertNotIn(
str(temp.id),
self.client_for("sec-staff").get("/api/uploads/").content.decode(),
)
def test_similarity_checks_are_private(self):
@@ -423,6 +443,36 @@ class ThrottleTests(SecurityTestCase):
{self.guest.get("/api/status/").status_code for _ in range(12)}, {200}
)
def test_signed_media_urls_are_not_throttled(self):
"""<img>/<video> tags fetch these without an Authorization header.
Regression: they were charged to the anonymous bucket, so galleries
and the fish-greeting download started returning 429 JSON instead of
the image bytes.
"""
item = self.make_item("throttle-media", owner=self.users["sec-uploader"])
codes = {
self.guest.get(f"/api/files/J-{item.id}/raw/").status_code
for _ in range(150)
}
self.assertEqual(codes, {200})
def test_staged_upload_files_are_not_throttled(self):
temp = TempUpload.objects.create(
user=self.users["sec-uploader"],
file=SimpleUploadedFile("throttle-temp.bin", b"staged"),
original_filename="throttle-temp.bin",
md5=hashlib.md5(b"throttle-temp").hexdigest(),
size=6,
)
signature = sign_payload(
{"temp": str(temp.id), "user": self.users["sec-uploader"].id},
services.UPLOAD_FILE_SALT,
)
url = f"/api/uploads/{temp.id}/file/?sig={signature}"
codes = {self.guest.get(url).status_code for _ in range(150)}
self.assertEqual(codes, {200})
class RemoteUrlTests(SecurityTestCase):
def test_allowlist(self):
+257 -34
View File
@@ -1,22 +1,39 @@
"""Minimal e621 API client for server-side matching and metadata refresh.
The SPA talks to e621 directly for browsing; this client exists for work the
browser cannot do reliably: long batch scans, and requests tied to a library
item rather than an open page. It uses the requesting user's stored
credentials and a global throttle (e621 asks for at most two requests per
second).
browser cannot do reliably: long batch scans, staged-upload processing and
requests tied to a library item rather than an open page. It uses the
requesting user's stored credentials and a global throttle (e621 asks for at
most two requests per second, one per second sustained).
e621's load balancer also sheds load with 429s (sometimes with an HTML
"shedding" page instead of JSON) and the IQDB endpoint has its own, much
stricter throttle. Every call therefore retries with exponential backoff and
honours ``Retry-After``; only 401/403 are treated as fatal.
"""
import logging
import random
import threading
import time
from pathlib import Path
import requests
from django.conf import settings
logger = logging.getLogger(__name__)
# Seconds between requests, per process. e621 allows 2/s hard and 1/s
# sustained; each gunicorn worker throttles on its own, so leave enough
# headroom that combined traffic does not trip the limit.
REQUEST_INTERVAL = 1.0
# IQDB is throttled far more aggressively than the rest of the API.
IQDB_INTERVAL = 2.0
MAX_ATTEMPTS = 4
BACKOFF_BASE = 2.0
BACKOFF_CAP = 60.0
RETRYABLE_STATUSES = {429, 500, 502, 503, 504}
class E621Error(Exception):
@@ -27,6 +44,14 @@ class E621NotFound(E621Error):
"""The requested post does not exist (HTTP 404)."""
class E621AuthError(E621Error):
"""e621 rejected the stored credentials (401/403)."""
class E621RateLimited(E621Error):
"""e621 shed load or throttled the request after every retry."""
_throttle_lock = threading.Lock()
_last_request_at = 0.0
@@ -35,56 +60,170 @@ def credentials_configured(user):
return bool(user is not None and getattr(user, "e621_configured", False))
def _wait_for_slot():
def _wait_for_slot(interval=REQUEST_INTERVAL):
global _last_request_at
with _throttle_lock:
delay = _last_request_at + REQUEST_INTERVAL - time.monotonic()
delay = _last_request_at + interval - time.monotonic()
if delay > 0:
time.sleep(delay)
_last_request_at = time.monotonic()
def _retry_delay(attempt, response=None):
"""Backoff for a retryable failure, honouring ``Retry-After``."""
if response is not None:
retry_after = response.headers.get("Retry-After")
if retry_after:
try:
return max(float(retry_after), 1.0)
except (TypeError, ValueError):
pass
delay = min(BACKOFF_BASE * (2**attempt), BACKOFF_CAP)
return delay + random.uniform(0, delay * 0.25)
def _request(
user,
method,
path,
*,
params=None,
data=None,
files=None,
timeout=30,
require_auth=True,
interval=REQUEST_INTERVAL,
attempts=MAX_ATTEMPTS,
):
"""One e621 call with retries; returns the parsed JSON payload.
``files`` may be a callable returning the multipart mapping, which is
called once per attempt: streamed uploads consume their file handle, so a
retry needs a freshly opened file.
Raises E621NotFound for 404s, E621AuthError for 401/403 and
E621RateLimited when e621 keeps shedding/throttling after every attempt.
"""
configured = credentials_configured(user)
if require_auth and not configured:
raise E621Error("Configure your e621 credentials in Account first.")
base = (getattr(user, "e621_base_url", "") or "https://e621.net").rstrip("/")
auth = (
(user.e621_username, user.e621_api_key_plain) if configured else None
)
url = f"{base}{path}"
last_error = None
for attempt in range(attempts):
request_files = files() if callable(files) else files
_wait_for_slot(interval)
response = None
try:
response = requests.request(
method,
url,
params=params,
data=data,
files=request_files,
auth=auth,
headers={"User-Agent": settings.USER_AGENT},
timeout=timeout,
)
except requests.RequestException as exc:
last_error = E621Error(f"Could not reach e621: {exc}")
else:
if response.status_code == 404:
raise E621NotFound(f"e621 returned 404 for {path}")
if response.status_code in {401, 403}:
raise E621AuthError(
f"e621 rejected the request ({response.status_code}). "
"Check the stored e621 credentials."
)
if response.status_code == 429:
# Throttles and load-shedding can arrive as JSON ({"message":
# "Throttled: ..."}) or as an HTML page.
message = ""
try:
payload = response.json()
except ValueError:
payload = None
if isinstance(payload, dict):
message = str(
payload.get("message") or payload.get("error") or ""
)
last_error = E621RateLimited(
message or f"e621 throttled the request for {path}"
)
elif response.status_code < 400:
try:
return response.json()
except ValueError:
# An HTML page with a 2xx status.
last_error = E621RateLimited(
f"e621 returned an unexpected {response.status_code} response."
)
elif response.status_code in RETRYABLE_STATUSES:
last_error = E621RateLimited(
f"e621 replied {response.status_code} for {path}"
)
else:
raise E621Error(f"e621 replied {response.status_code} for {path}")
finally:
_close_upload_files(request_files)
if attempt + 1 < attempts:
delay = _retry_delay(attempt, response)
logger.info(
"e621 %s %s failed (%s); retrying in %.1fs",
method,
path,
last_error,
delay,
)
time.sleep(delay)
if last_error is None:
last_error = E621Error("e621 request failed.")
raise last_error
def _close_upload_files(files):
"""Close the handles behind a multipart mapping (see _request)."""
if not isinstance(files, dict):
return
for value in files.values():
handle = value[1] if isinstance(value, tuple) and len(value) > 1 else value
close = getattr(handle, "close", None)
if close is not None:
try:
close()
except Exception: # noqa: BLE001 - closing must never mask errors
pass
def get(user, path, params=None, timeout=30, require_auth=True):
"""GET an e621 API path using the user's credentials.
Reads that e621 serves anonymously (searches, pools, tags) can pass
require_auth=False; matching endpoints keep requiring credentials.
Raises E621NotFound for 404s and E621Error for everything else that isn't
a 2xx, so callers never see requests exceptions.
"""
configured = credentials_configured(user)
if require_auth and not configured:
raise E621Error("Configure your e621 credentials in Account first.")
base = (getattr(user, "e621_base_url", "") or "https://e621.net").rstrip("/")
_wait_for_slot()
try:
response = requests.get(
f"{base}{path}",
return _request(
user,
"GET",
path,
params=params,
auth=(
(user.e621_username, user.e621_api_key_plain)
if configured
else None
),
headers={"User-Agent": settings.USER_AGENT},
timeout=timeout,
require_auth=require_auth,
)
except requests.RequestException as exc:
raise E621Error(f"Could not reach e621: {exc}") from exc
if response.status_code == 404:
raise E621NotFound(f"e621 returned 404 for {path}")
if response.status_code >= 400:
raise E621Error(f"e621 replied {response.status_code} for {path}")
try:
return response.json()
except ValueError as exc:
raise E621Error("e621 returned an unexpected response.") from exc
def find_post_by_md5(user, md5):
"""The e621 post with this exact MD5, or None."""
payload = get(user, "/posts.json", params={"tags": f"md5:{md5}", "limit": 1})
payload = get(
user,
"/posts.json",
params={"tags": f"md5:{md5}", "limit": 1},
require_auth=False,
)
posts = payload.get("posts") if isinstance(payload, dict) else None
if not posts:
return None
@@ -98,3 +237,87 @@ def fetch_post(user, post_id):
if not isinstance(post, dict):
raise E621Error("e621 returned an unexpected post payload.")
return post
def check_md5_batch(user, md5s):
"""Look many MD5s up in one posts.json query.
Returns ``{md5: post}`` for the ones e621 knows; missing MD5s are simply
absent. Works anonymously, like the original app's batch cache command.
"""
wanted = {str(value).strip().lower() for value in md5s if value}
if not wanted:
return {}
values = sorted(wanted)
payload = get(
user,
"/posts.json",
params={
"tags": f"md5:{','.join(values)}",
"limit": min(len(values), 320),
},
require_auth=False,
)
posts = payload.get("posts") if isinstance(payload, dict) else None
found = {}
for post in posts or []:
if not isinstance(post, dict):
continue
file_data = post.get("file") or {}
md5 = str(file_data.get("md5") or "").strip().lower()
if md5 in wanted:
found[md5] = post
return found
def fetch_posts_by_ids(user, ids):
"""Fetch many posts in one query (up to 320 ids). Missing ids are absent."""
values = sorted({int(value) for value in ids})
if not values:
return []
payload = get(
user,
"/posts.json",
params={
"tags": f"id:{','.join(str(value) for value in values)}",
"limit": min(len(values), 320),
},
require_auth=False,
)
posts = payload.get("posts") if isinstance(payload, dict) else None
return [post for post in posts or [] if isinstance(post, dict)]
def iqdb_search(user, path, timeout=60):
"""Reverse-image search one file through e621's IQDB endpoint.
Returns the legacy match list. Uses the extra-strict IQDB interval and
retries through e621's throttle; raises E621RateLimited when it persists.
"""
path = Path(path)
def open_file():
# A fresh handle per attempt: the stream is consumed by the request.
return {"search[file]": (path.name, open(path, "rb"))}
payload = _request(
user,
"POST",
"/iqdb_queries.json",
files=open_file,
timeout=timeout,
require_auth=False,
interval=IQDB_INTERVAL,
)
if isinstance(payload, list):
return payload
if isinstance(payload, dict):
matches = payload.get("matches")
if isinstance(matches, list):
return matches
# e621 answers its throttle with {"success": false, "message": ...}.
message = payload.get("message") or payload.get("error")
if message:
raise E621RateLimited(str(message))
raise E621Error("e621 returned an unexpected IQDB payload.")
return []
@@ -18,6 +18,9 @@ class Command(BaseCommand):
)
def handle(self, *args, **options):
from apps.library.upload_pipeline import reap_stale_claims
reap_stale_claims()
cutoff = timezone.now() - timedelta(hours=options["hours"])
queryset = TempUpload.objects.filter(created_at__lt=cutoff)
if not options["include_completed"]:
@@ -0,0 +1,68 @@
"""Process staged uploads: e621 MD5, visual similarity, IQDB.
Runs synchronously, unlike the daemon thread the API starts on demand. Useful
for tests, for a manual drain after an outage and for the scheduler if a
deployment wants a periodic safety net.
python manage.py process_uploads # every user with queued work, once
python manage.py process_uploads --user 3 # one user
python manage.py process_uploads --loop 60 # keep draining every 60s
"""
import time
from django.core.management.base import BaseCommand
from apps.library.models import TempUpload
from apps.library.upload_pipeline import (
MAX_ATTEMPTS,
OUTSTANDING_Q,
WORK_STATUSES,
reap_stale_claims,
run_pipeline,
)
class Command(BaseCommand):
help = "Run the staged-upload pipeline (MD5 -> visual similarity -> IQDB)."
def add_arguments(self, parser):
parser.add_argument(
"--user",
type=int,
default=None,
help="Only process this user id.",
)
parser.add_argument(
"--loop",
type=int,
default=0,
metavar="SECONDS",
help="Keep draining every SECONDS seconds instead of exiting.",
)
def handle(self, *args, **options):
interval = options["loop"] or 0
while True:
self.drain(user_id=options["user"])
if interval <= 0:
return
time.sleep(interval)
def drain(self, user_id=None):
reap_stale_claims()
queryset = (
TempUpload.objects.filter(status__in=WORK_STATUSES)
.filter(OUTSTANDING_Q)
.filter(attempts__lt=MAX_ATTEMPTS)
)
if user_id is not None:
queryset = queryset.filter(user_id=user_id)
user_ids = list(queryset.values_list("user_id", flat=True).distinct())
if not user_ids:
self.stdout.write("No staged uploads need processing.")
return
for value in user_ids:
self.stdout.write(f"Processing staged uploads for user {value}...")
run_pipeline(value)
self.stdout.write(self.style.SUCCESS(f"Processed {len(user_ids)} queue(s)."))
@@ -0,0 +1,81 @@
# Generated by Django 6.1.1 on 2026-09-21 13:27
import django.db.models.deletion
from django.conf import settings
from django.db import migrations, models
def backfill_pipeline_checks(apps, schema_editor):
"""Mark pre-pipeline rows as already MD5/visual-checked when they were.
Rows with an e621 post id (or already completed) clearly went through the
MD5 phase; rows with visual matches went through the visual phase.
Everything else stays unset so the new pipeline picks it up once after
deploy — a re-check of stale pending uploads is the desired behavior.
"""
TempUpload = apps.get_model("library", "TempUpload")
TempUpload.objects.filter(
models.Q(e621_post_id__isnull=False) | models.Q(status="completed")
).update(e621_checked_at=models.F("updated_at"))
TempUpload.objects.filter(visual_matches__isnull=False).update(
visual_checked_at=models.F("updated_at")
)
class Migration(migrations.Migration):
dependencies = [
('library', '0010_similaritycheck'),
migrations.swappable_dependency(settings.AUTH_USER_MODEL),
]
operations = [
migrations.AddField(
model_name='tempupload',
name='attempts',
field=models.PositiveSmallIntegerField(default=0),
),
migrations.AddField(
model_name='tempupload',
name='claimed_at',
field=models.DateTimeField(blank=True, db_index=True, null=True),
),
migrations.AddField(
model_name='tempupload',
name='e621_checked_at',
field=models.DateTimeField(blank=True, null=True),
),
migrations.AddField(
model_name='tempupload',
name='pipeline_error',
field=models.TextField(blank=True, default=''),
),
migrations.AddField(
model_name='tempupload',
name='visual_checked_at',
field=models.DateTimeField(blank=True, null=True),
),
migrations.CreateModel(
name='UploadRun',
fields=[
('id', models.BigAutoField(auto_created=True, primary_key=True, serialize=False, verbose_name='ID')),
('status', models.CharField(choices=[('idle', 'Idle'), ('running', 'Running'), ('paused', 'Paused'), ('error', 'Error')], default='idle', max_length=20)),
('phase', models.CharField(blank=True, choices=[('', 'None'), ('md5', 'e621 MD5'), ('visual', 'Visual similarity'), ('iqdb', 'IQDB')], default='', max_length=20)),
('total', models.IntegerField(default=0)),
('processed', models.IntegerField(default=0)),
('matched', models.IntegerField(default=0)),
('failed', models.IntegerField(default=0)),
('error', models.TextField(blank=True, default='')),
('started_at', models.DateTimeField(blank=True, null=True)),
('updated_at', models.DateTimeField(auto_now=True)),
('user', models.OneToOneField(on_delete=django.db.models.deletion.CASCADE, related_name='upload_run', to=settings.AUTH_USER_MODEL)),
],
options={
'ordering': ['-updated_at'],
},
),
migrations.RunPython(
code=backfill_pipeline_checks, reverse_code=migrations.RunPython.noop
),
]
@@ -0,0 +1,34 @@
# Generated by Django 6.1.1 on 2026-09-23
from django.db import migrations, models
def purge_completed_uploads(apps, schema_editor):
"""Completed staged uploads are notifications, not records.
The board used to keep every auto-uploaded/indexed row until it was
dismissed by hand, so they accumulated without bound (and bulk dismissal
is capped at 1000 ids). Completions now live in a small per-run feed, so
the old rows are removed here.
"""
TempUpload = apps.get_model("library", "TempUpload")
for temp in TempUpload.objects.filter(status="completed").iterator():
if temp.file:
temp.file.delete(save=False)
temp.delete()
class Migration(migrations.Migration):
dependencies = [
("library", "0011_upload_pipeline"),
]
operations = [
migrations.AddField(
model_name="uploadrun",
name="recent_completions",
field=models.JSONField(blank=True, default=list),
),
migrations.RunPython(purge_completed_uploads, migrations.RunPython.noop),
]
+70
View File
@@ -164,6 +164,16 @@ class TempUpload(models.Model):
on_delete=models.SET_NULL,
related_name="temp_uploads",
)
# Background pipeline bookkeeping (see apps/library/upload_pipeline.py).
# e621_checked_at/visual_checked_at are set once the corresponding phase
# ran, so "never checked" and "checked, nothing found" stay distinct.
e621_checked_at = models.DateTimeField(null=True, blank=True)
visual_checked_at = models.DateTimeField(null=True, blank=True)
# Worker claim for cross-process mutual exclusion; stale claims are
# reaped and the row queued again.
claimed_at = models.DateTimeField(null=True, blank=True, db_index=True)
pipeline_error = models.TextField(blank=True, default="")
attempts = models.PositiveSmallIntegerField(default=0)
created_at = models.DateTimeField(auto_now_add=True)
updated_at = models.DateTimeField(auto_now=True)
@@ -174,6 +184,66 @@ class TempUpload(models.Model):
return f"{self.original_filename} ({self.status})"
class UploadRun(models.Model):
"""Per-user state of the background upload pipeline.
One row per user acts as the cheap status source the SPA polls and as a
place for batch-level failures (broken e621 credentials, outages) that
would otherwise be repeated on every row.
"""
STATUS_IDLE = "idle"
STATUS_RUNNING = "running"
STATUS_PAUSED = "paused"
STATUS_ERROR = "error"
STATUS_CHOICES = [
(STATUS_IDLE, "Idle"),
(STATUS_RUNNING, "Running"),
(STATUS_PAUSED, "Paused"),
(STATUS_ERROR, "Error"),
]
PHASE_MD5 = "md5"
PHASE_VISUAL = "visual"
PHASE_IQDB = "iqdb"
PHASE_CHOICES = [
("", "None"),
(PHASE_MD5, "e621 MD5"),
(PHASE_VISUAL, "Visual similarity"),
(PHASE_IQDB, "IQDB"),
]
user = models.OneToOneField(
settings.AUTH_USER_MODEL,
on_delete=models.CASCADE,
related_name="upload_run",
)
status = models.CharField(
max_length=20, choices=STATUS_CHOICES, default=STATUS_IDLE
)
phase = models.CharField(
max_length=20, choices=PHASE_CHOICES, blank=True, default=""
)
total = models.IntegerField(default=0)
processed = models.IntegerField(default=0)
matched = models.IntegerField(default=0)
failed = models.IntegerField(default=0)
error = models.TextField(blank=True, default="")
# Rolling feed of recent completions for the upload board. Completed
# staged uploads are deleted as soon as they are indexed; this only tells
# the live page "filename -> J-x" while it watches. Bounded, never
# dismissed, and ignored by fresh page loads.
recent_completions = models.JSONField(default=list, blank=True)
started_at = models.DateTimeField(null=True, blank=True)
updated_at = models.DateTimeField(auto_now=True)
class Meta:
ordering = ["-updated_at"]
def __str__(self):
return f"Upload run for {self.user_id} ({self.status})"
class DownloadTask(models.Model):
"""A background 'Download to Library' job with progress tracking."""
+101 -5
View File
@@ -3,7 +3,6 @@ from datetime import timedelta
from pathlib import Path
from django.conf import settings
from django.core import signing
from rest_framework import serializers
from .models import (
@@ -20,6 +19,7 @@ from .services import (
VIDEO_EXTENSIONS,
signed_media_url,
)
from .signing_urls import sign_payload
class MediaLocationSerializer(serializers.ModelSerializer):
@@ -145,6 +145,12 @@ class TempUploadSerializer(serializers.ModelSerializer):
library_j_id = serializers.SerializerMethodField()
file_url = serializers.SerializerMethodField()
preview_url = serializers.SerializerMethodField()
md5_checked = serializers.SerializerMethodField()
visual_checked = serializers.SerializerMethodField()
iqdb_checked = serializers.SerializerMethodField()
processing = serializers.SerializerMethodField()
similar_count = serializers.SerializerMethodField()
visual_matches = serializers.SerializerMethodField()
class Meta:
model = TempUpload
@@ -165,6 +171,13 @@ class TempUploadSerializer(serializers.ModelSerializer):
"library_j_id",
"file_url",
"preview_url",
"pipeline_error",
"attempts",
"md5_checked",
"visual_checked",
"iqdb_checked",
"processing",
"similar_count",
"created_at",
"updated_at",
]
@@ -187,9 +200,9 @@ class TempUploadSerializer(serializers.ModelSerializer):
user = self._request_user()
if user is None:
return None
signature = signing.dumps(
signature = sign_payload(
{"temp": str(obj.id), "user": user.id},
salt=UPLOAD_FILE_SALT,
UPLOAD_FILE_SALT,
)
url = f"/api/uploads/{obj.id}/file/?sig={signature}"
request = self.context.get("request")
@@ -215,6 +228,89 @@ class TempUploadSerializer(serializers.ModelSerializer):
item, user, action, request=self.context.get("request")
)
def get_md5_checked(self, obj):
return obj.e621_checked_at is not None
def get_visual_checked(self, obj):
return obj.visual_checked_at is not None
def get_iqdb_checked(self, obj):
return obj.iqdb_data is not None
def get_processing(self, obj):
return obj.claimed_at is not None
def get_similar_count(self, obj):
return len(obj.iqdb_data or []) + len(obj.visual_matches or [])
@staticmethod
def _visual_item_id(entry):
if not isinstance(entry, dict):
return None
item_id = entry.get("item_id")
if item_id is None:
j_id = str(entry.get("j_id") or "")
if j_id.upper().startswith("J-"):
j_id = j_id[2:]
item_id = j_id if j_id.isdigit() else None
try:
return int(item_id)
except (TypeError, ValueError):
return None
def get_visual_matches(self, obj):
"""Rebuild match rows with fresh signed thumbnail URLs.
Storing the signed URL meant it aged out (or came from an older
signing scheme) and the "Already in your library" grid showed broken
tiles. The stored rows only carry the item reference now.
"""
entries = obj.visual_matches or []
if not entries:
return entries
wanted = {}
for entry in entries:
item_id = self._visual_item_id(entry)
if item_id is not None:
wanted[item_id] = None
items = MediaItem.objects.in_bulk(list(wanted))
user = self._request_user()
request = self.context.get("request")
matches = []
for entry in entries:
item_id = self._visual_item_id(entry)
item = items.get(item_id) if item_id is not None else None
if item is None:
continue
matches.append(
{
"j_id": f"J-{item.id}",
"filename": entry.get("filename") or item.md5,
"similarity": entry.get("similarity"),
"thumbnail_url": signed_media_url(
item, user, "thumbnail", request=request
),
}
)
return matches
class TempUploadListSerializer(TempUploadSerializer):
"""Compact staged-upload row for the board and the status polling.
Drops the heavy post/IQDB payloads (the metadata modal fetches the full
row) while keeping the pipeline flags the board renders per tile.
"""
class Meta(TempUploadSerializer.Meta):
fields = [
field
for field in TempUploadSerializer.Meta.fields
if field
not in {"e621_data", "iqdb_data", "visual_matches", "custom_tags", "custom_notes"}
]
read_only_fields = fields
class DownloadTaskSerializer(serializers.ModelSerializer):
task_id = serializers.UUIDField(source="id", read_only=True)
@@ -295,9 +391,9 @@ class SimilarityCheckSerializer(serializers.ModelSerializer):
user = self._request_user()
if user is None or not obj.file:
return None
signature = signing.dumps(
signature = sign_payload(
{"check": str(obj.id), "user": user.id},
salt=UPLOAD_FILE_SALT,
UPLOAD_FILE_SALT,
)
url = f"/api/similarity/{obj.id}/file/?sig={signature}"
request = self.context.get("request")
+137 -15
View File
@@ -6,16 +6,21 @@ import os
import re
import shutil
import subprocess
import uuid
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlencode
import imagehash
from django.conf import settings
from django.core import signing
from django.http import FileResponse, Http404, HttpResponse
from django.utils.cache import get_conditional_response
from django.utils.http import http_date
from django.utils.text import get_valid_filename
from PIL import Image
from PIL import Image, ImageOps
from .models import MediaItem, MediaLocation
from .signing_urls import sign_payload
logger = logging.getLogger(__name__)
@@ -37,6 +42,11 @@ UPLOAD_FILE_SALT = "j621.upload-file"
MEDIA_FILE_SALT = "j621.media-file"
CHUNK_SIZE = 1024 * 1024
RANGE_RE = re.compile(r"bytes=(\d*)-(\d*)$")
# Versioned media URLs are immutable, so they may sit in the browser cache for
# as long as the signature is guaranteed to stay valid (7 days).
MEDIA_CACHE_SECONDS = 6 * 86400
# Staged uploads and similarity files can be deleted at any moment.
TEMP_CACHE_SECONDS = 3600
def compute_md5(path):
@@ -78,14 +88,24 @@ def signed_media_url(item, user, action="raw", request=None):
With a ``request`` the URL is absolute, so the SPA also works when it is
served from a different origin; without one it stays relative.
The ``v`` parameter is the item's MD5: it busts the browser cache exactly
when the file is replaced (the optimize flow rewrites files under the same
J-ID), which is what lets the URL be cached for days instead of re-minted
on every response.
"""
path = f"/api/files/J-{item.id}/{action}/"
params = {}
if user is not None and getattr(user, "is_authenticated", False):
signature = signing.dumps(
params = {
"v": item.md5,
"sig": sign_payload(
{"item": item.id, "user": user.id, "action": action},
salt=MEDIA_FILE_SALT,
)
path = f"{path}?sig={signature}"
MEDIA_FILE_SALT,
),
}
if params:
path = f"{path}?{urlencode(params)}"
if request is None:
return path
return request.build_absolute_uri(path)
@@ -207,12 +227,36 @@ class RangeFileWrapper:
self.file.close()
def serve_file(request, path, download=False):
"""Serve a file with HTTP range support (needed for video seeking)."""
def _apply_cache_headers(response, cache_control, etag, mtime):
response["Cache-Control"] = cache_control
response["ETag"] = etag
response["Last-Modified"] = http_date(mtime)
return response
def serve_file(request, path, download=False, *, max_age=TEMP_CACHE_SECONDS, immutable=False):
"""Serve a file with HTTP range support (needed for video seeking).
Responses carry validators (ETag/Last-Modified) and a private
``Cache-Control`` so browsers reuse media instead of re-downloading it on
every SPA poll. ``max_age``/``immutable`` are chosen by the caller: versioned
library media can be cached hard, staged files only briefly.
"""
path = Path(path)
if not path.is_file():
raise Http404
size = path.stat().st_size
stat = path.stat()
size = stat.st_size
etag = f'W/"{size:x}-{stat.st_mtime_ns:x}"'
last_modified = datetime.fromtimestamp(stat.st_mtime, tz=timezone.utc)
conditional = get_conditional_response(
request, etag=etag, last_modified=last_modified
)
if conditional is not None:
return conditional
cache_control = f"private, max-age={int(max_age)}"
if immutable:
cache_control += ", immutable"
content_type = mimetypes.guess_type(str(path))[0] or "application/octet-stream"
range_header = request.headers.get("Range", "").strip()
if range_header:
@@ -240,7 +284,9 @@ def serve_file(request, path, download=False):
response["Content-Length"] = str(length)
response["Content-Range"] = f"bytes {start}-{end}/{size}"
response["Accept-Ranges"] = "bytes"
return response
return _apply_cache_headers(
response, cache_control, etag, stat.st_mtime
)
response = FileResponse(
open(path, "rb"),
content_type=content_type,
@@ -248,7 +294,7 @@ def serve_file(request, path, download=False):
filename=path.name,
)
response["Accept-Ranges"] = "bytes"
return response
return _apply_cache_headers(response, cache_control, etag, stat.st_mtime)
class DownloadCancelled(Exception):
@@ -436,15 +482,38 @@ def sanitize_iqdb_results(results):
return cleaned
def _thumbnail_is_fresh(target, path):
"""True when the cached thumbnail exists and is at least as new as source."""
try:
stat = target.stat()
if stat.st_size <= 0:
return False
return stat.st_mtime >= os.path.getmtime(path)
except OSError:
return False
def _thumbs_dir():
thumbs_dir = Path(settings.MEDIA_ROOT) / "thumbs"
thumbs_dir.mkdir(parents=True, exist_ok=True)
return thumbs_dir
def generate_video_thumbnail(md5, path):
"""Extract a JPEG thumbnail from a video, cached under MEDIA_ROOT/thumbs."""
if not shutil.which("ffmpeg"):
return None
thumbs_dir = Path(settings.MEDIA_ROOT) / "thumbs"
thumbs_dir.mkdir(parents=True, exist_ok=True)
try:
thumbs_dir = _thumbs_dir()
except OSError:
logger.exception("Could not create the thumbnail folder")
return None
target = thumbs_dir / f"{md5}.jpg"
if target.exists() and target.stat().st_mtime >= os.path.getmtime(path):
if _thumbnail_is_fresh(target, path):
return target
# Write beside the target and move it into place, so a concurrent request
# can never read a half-written JPEG.
temp = thumbs_dir / f".{md5}.{uuid.uuid4().hex}.part.jpg"
command = [
"ffmpeg",
"-y",
@@ -458,10 +527,63 @@ def generate_video_thumbnail(md5, path):
"scale=480:-2",
"-loglevel",
"error",
str(target),
str(temp),
]
try:
subprocess.run(command, check=True, capture_output=True, timeout=60)
os.replace(temp, target)
except (subprocess.SubprocessError, OSError):
logger.exception("Could not build a video thumbnail for %s", path)
temp.unlink(missing_ok=True)
return None
return target if target.exists() else None
def generate_image_thumbnail(md5, path):
"""Downscale an image, cached under MEDIA_ROOT/thumbs like video thumbs.
The thumbnail action used to serve full-size originals for images; a
cached 480px JPEG keeps the library grid light without touching the
original file. Returns ``None`` when the source is missing or Pillow
cannot decode it, so callers can fall back to the original.
"""
try:
thumbs_dir = _thumbs_dir()
except OSError:
logger.exception("Could not create the thumbnail folder")
return None
target = thumbs_dir / f"{md5}.jpg"
if _thumbnail_is_fresh(target, path):
return target
temp = thumbs_dir / f".{md5}.{uuid.uuid4().hex}.part.jpg"
try:
with Image.open(path) as image:
# Animated formats: the first frame is the preview.
image.seek(0)
frame = ImageOps.exif_transpose(image) or image
frame = frame.convert("RGB")
frame.thumbnail((480, 480))
frame.save(temp, "JPEG", quality=82, optimize=True)
os.replace(temp, target)
except Exception: # noqa: BLE001 - previews must never break serving
logger.exception("Could not build an image thumbnail for %s", path)
temp.unlink(missing_ok=True)
return None
return target if target.exists() else None
def ensure_thumbnail(item):
"""Generate an item's cached thumbnail if it is missing or stale.
Warming thumbnails when a file is indexed keeps image decoding out of the
request path, where the upload pipeline's hashing used to starve it.
"""
location = item.locations.first()
if location is None:
return None
path = Path(location.path)
if not path.is_file():
return None
if path.suffix.lower() in VIDEO_EXTENSIONS:
return generate_video_thumbnail(item.md5, path)
return generate_image_thumbnail(item.md5, path)
+64
View File
@@ -0,0 +1,64 @@
"""Stable, expiring signatures for media URLs.
The SPA loads media with ``<img>``/``<video>`` tags, which cannot send the
API's ``Authorization`` header, so those URLs carry a signature instead. The
signature has to be *stable*: a URL that changes on every response makes the
browser treat every refetch as a new resource and re-download the file.
URLs are signed with a plain ``Signer`` (no per-second timestamp) plus an
explicit ``exp`` claim quantized to a bucket, so every request inside a bucket
mints the exact same URL. The URL rotates once per bucket and is valid for at
least ``URL_TTL_SECONDS`` and at most ``URL_TTL_SECONDS + URL_BUCKET_SECONDS``.
"""
import time
from django.core import signing
URL_TTL_SECONDS = 7 * 86400
URL_BUCKET_SECONDS = 24 * 3600
_BUCKETS = URL_TTL_SECONDS // URL_BUCKET_SECONDS
def _expiry(now=None):
current = time.time() if now is None else now
bucket = int(current // URL_BUCKET_SECONDS)
return (bucket + _BUCKETS + 1) * URL_BUCKET_SECONDS
def sign_payload(payload, salt, now=None):
"""Sign a payload with a stable, bucket-quantized expiry."""
return signing.Signer(salt=salt).sign_object(
{**payload, "exp": _expiry(now)}
)
def load_payload(signature, salt, legacy_max_age=86400):
"""Verify a signed payload; ``None`` when missing, tampered with or expired.
Legacy ``TimestampSigner`` values are still accepted for one release.
Detect them by their extra separator (``payload:timestamp:signature``):
a plain ``Signer`` accepts the HMAC a ``TimestampSigner`` computed over
``payload:timestamp`` and then chokes on the embedded timestamp while
decoding the JSON payload, which used to surface as a 500.
"""
if not signature:
return None
if signature.count(":") >= 2:
try:
return signing.TimestampSigner(salt=salt).unsign_object(
signature, max_age=legacy_max_age
)
except (signing.BadSignature, ValueError):
return None
try:
data = signing.Signer(salt=salt).unsign_object(signature)
except (signing.BadSignature, ValueError):
return None
if not isinstance(data, dict):
return None
try:
expired = int(data.get("exp", 0)) < time.time()
except (TypeError, ValueError):
return None
return None if expired else data
+8 -10
View File
@@ -11,7 +11,6 @@ from datetime import timedelta
from pathlib import Path
from django.conf import settings
from django.core import signing
from django.utils import timezone
from rest_framework import mixins, status, viewsets
from rest_framework.decorators import action
@@ -22,6 +21,7 @@ from rest_framework.response import Response
from . import services
from .models import MediaItem, SimilarityCheck
from .serializers import SimilarityCheckSerializer
from .signing_urls import load_payload
from .tools import item_brief
from .uploads import find_library_matches
@@ -136,7 +136,12 @@ class SimilarityCheckViewSet(
self.get_serializer(check).data, status=status.HTTP_201_CREATED
)
@action(detail=True, methods=["get", "head"], permission_classes=[AllowAny])
@action(
detail=True,
methods=["get", "head"],
permission_classes=[AllowAny],
throttle_classes=[],
)
def file(self, request, pk=None):
"""Serve the temp file; accepts a signed URL like staged uploads."""
check = None
@@ -144,14 +149,7 @@ class SimilarityCheckViewSet(
check = self.get_queryset().filter(pk=pk).first()
else:
signature = request.query_params.get("sig")
payload = None
if signature:
try:
payload = signing.loads(
signature, salt=services.UPLOAD_FILE_SALT, max_age=86400
)
except signing.BadSignature:
payload = None
payload = load_payload(signature, services.UPLOAD_FILE_SALT) if signature else None
if payload and payload.get("check") == str(pk):
check = SimilarityCheck.objects.filter(pk=pk).first()
if check is None or not check.file:
+99
View File
@@ -0,0 +1,99 @@
"""The server-side e621 client: batching, retries and IQDB stream handling."""
from pathlib import Path
from tempfile import TemporaryDirectory
from unittest import mock
from django.test import SimpleTestCase
from apps.library import e621
class FakeResponse:
def __init__(self, status_code, payload=None, headers=None):
self.status_code = status_code
self._payload = payload
self.headers = headers or {}
def json(self):
if self._payload is None:
raise ValueError("not json")
return self._payload
class E621ClientTests(SimpleTestCase):
def test_iqdb_search_reopens_the_file_on_retry(self):
bodies = []
def fake_request(method, url, **kwargs):
handle = kwargs["files"]["search[file]"][1]
bodies.append(handle.read())
if len(bodies) == 1:
return FakeResponse(
429, {"success": False, "message": "Throttled"}
)
return FakeResponse(200, [{"post_id": 1, "score": 90.0}])
with TemporaryDirectory() as tmp:
path = Path(tmp) / "x.png"
path.write_bytes(b"image-bytes")
with mock.patch.object(
e621.requests, "request", side_effect=fake_request
), mock.patch.object(e621.time, "sleep"), mock.patch.object(
e621, "_wait_for_slot"
):
result = e621.iqdb_search(None, path)
# Both attempts must send the full body, not the consumed handle.
self.assertEqual(bodies, [b"image-bytes", b"image-bytes"])
self.assertEqual(result, [{"post_id": 1, "score": 90.0}])
def test_check_md5_batch_keys_by_md5(self):
payload = {
"posts": [
{"id": 5, "file": {"md5": "a" * 32}},
{"id": 6, "file": {"md5": "b" * 32}},
]
}
with mock.patch.object(e621, "get", return_value=payload) as getter:
found = e621.check_md5_batch(None, ["A" * 32, "b" * 32])
self.assertEqual(set(found), {"a" * 32, "b" * 32})
params = getter.call_args.kwargs["params"]
self.assertTrue(params["tags"].startswith("md5:"))
self.assertEqual(params["limit"], 2)
def test_auth_errors_are_not_retried(self):
calls = []
def fake_request(*args, **kwargs):
calls.append(1)
return FakeResponse(403, {"error": "nope"})
with mock.patch.object(
e621.requests, "request", side_effect=fake_request
), mock.patch.object(e621, "_wait_for_slot"):
with self.assertRaises(e621.E621AuthError):
e621._request(None, "GET", "/posts.json", require_auth=False)
self.assertEqual(len(calls), 1)
def test_load_shedding_html_raises_rate_limited_after_retries(self):
with mock.patch.object(
e621.requests,
"request",
return_value=FakeResponse(200, None),
), mock.patch.object(e621.time, "sleep"), mock.patch.object(
e621, "_wait_for_slot"
):
with self.assertRaises(e621.E621RateLimited):
e621._request(
None, "GET", "/posts.json", require_auth=False, attempts=2
)
def test_fetch_posts_by_ids_queries_with_id_tag(self):
with mock.patch.object(
e621, "get", return_value={"posts": [{"id": 9}]}
) as getter:
posts = e621.fetch_posts_by_ids(None, [9])
self.assertEqual(posts, [{"id": 9}])
params = getter.call_args.kwargs["params"]
self.assertEqual(params["tags"], "id:9")
@@ -0,0 +1,188 @@
"""Signed media URLs must be stable, versioned and cacheable.
Regression: signatures embedded the current second, so every API response
re-minted every URL and browsers re-downloaded each image on every poll; the
file responses also carried no cache headers at all.
"""
import base64
import hashlib
import io
import shutil
import tempfile
import time
from pathlib import Path
from unittest import mock
from django.contrib.auth import get_user_model
from django.core import signing
from django.core.files.uploadedfile import SimpleUploadedFile
from django.test import Client, TestCase, override_settings
from PIL import Image
from rest_framework.authtoken.models import Token
from apps.library import services
from apps.library.models import MediaItem, MediaLocation, TempUpload
from apps.library.signing_urls import load_payload, sign_payload
User = get_user_model()
TINY_PNG = base64.b64decode(
"iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAYAAAAfFcSJAAAADUlEQVR42mP8z8BQDwAEhQGAhKmMIQAAAABJRU5ErkJggg=="
)
def png_bytes(width=1200, height=800, color=(20, 120, 200)):
buffer = io.BytesIO()
Image.new("RGB", (width, height), color).save(buffer, format="PNG")
return buffer.getvalue()
class MediaCacheTests(TestCase):
@classmethod
def setUpClass(cls):
super().setUpClass()
cls._tmp = tempfile.mkdtemp(prefix="j621-cache-")
cls._media = Path(cls._tmp) / "media"
cls._watched = cls._media / "library"
cls._watched.mkdir(parents=True, exist_ok=True)
cls._settings = override_settings(
MEDIA_ROOT=str(cls._media), WATCHED_FOLDER=str(cls._watched)
)
cls._settings.enable()
@classmethod
def tearDownClass(cls):
cls._settings.disable()
shutil.rmtree(cls._tmp, ignore_errors=True)
super().tearDownClass()
def setUp(self):
self.user = User.objects.create_user(
username="cache-uploader", password="cache-pass-123456"
)
self.user.role = "uploader"
self.user.save(update_fields=["role"])
self.token = Token.objects.create(user=self.user).key
payload = png_bytes()
path = self._watched / "cache-image.png"
path.write_bytes(payload)
self.item = MediaItem.objects.create(
md5=hashlib.md5(payload).hexdigest(), size=len(payload)
)
MediaLocation.objects.create(
item=self.item, path=str(path), rel_path=path.name, mtime=time.time()
)
self.client = Client()
def signed(self, action):
return {
"sig": sign_payload(
{"item": self.item.id, "user": self.user.id, "action": action},
services.MEDIA_FILE_SALT,
)
}
def test_media_response_carries_cache_headers(self):
response = self.client.get(
f"/api/files/J-{self.item.id}/raw/", self.signed("raw")
)
self.assertEqual(response.status_code, 200)
self.assertIn("private", response["Cache-Control"])
self.assertIn(
f"max-age={services.MEDIA_CACHE_SECONDS}", response["Cache-Control"]
)
self.assertIn("immutable", response["Cache-Control"])
self.assertTrue(response["ETag"])
self.assertTrue(response["Last-Modified"])
def test_media_revalidation_returns_304(self):
url = f"/api/files/J-{self.item.id}/raw/"
first = self.client.get(url, self.signed("raw"))
second = self.client.get(
url, self.signed("raw"), HTTP_IF_NONE_MATCH=first["ETag"]
)
self.assertEqual(second.status_code, 304)
self.assertEqual(second.content, b"")
def test_image_thumbnail_is_generated_and_reused(self):
url = f"/api/files/J-{self.item.id}/thumbnail/"
response = self.client.get(url, self.signed("thumbnail"))
self.assertEqual(response.status_code, 200)
self.assertEqual(response["Content-Type"], "image/jpeg")
thumb = self._media / "thumbs" / f"{self.item.md5}.jpg"
self.assertTrue(thumb.exists())
with Image.open(thumb) as image:
self.assertLessEqual(max(image.size), 480)
before = thumb.stat().st_mtime_ns
self.client.get(url, self.signed("thumbnail"))
self.assertEqual(thumb.stat().st_mtime_ns, before)
def test_ensure_thumbnail_reuses_the_cache(self):
first = services.ensure_thumbnail(self.item)
self.assertIsNotNone(first)
self.assertTrue(first.exists())
mtime = first.stat().st_mtime_ns
second = services.ensure_thumbnail(self.item)
self.assertEqual(second, first)
self.assertEqual(second.stat().st_mtime_ns, mtime)
def test_thumbnail_of_a_missing_source_does_not_error(self):
"""Regression: getmtime() on a vanished source used to raise a 500."""
thumbs = self._media / "thumbs"
thumbs.mkdir(parents=True, exist_ok=True)
(thumbs / f"{self.item.md5}.jpg").write_bytes(b"stale")
Path(self.item.locations.first().path).unlink()
self.assertIsNone(services.ensure_thumbnail(self.item))
response = self.client.get(
f"/api/files/J-{self.item.id}/thumbnail/", self.signed("thumbnail")
)
self.assertEqual(response.status_code, 404)
def test_staged_files_cache_briefly(self):
temp = TempUpload.objects.create(
user=self.user,
file=SimpleUploadedFile("staged.png", TINY_PNG, content_type="image/png"),
original_filename="staged.png",
md5=hashlib.md5(b"staged").hexdigest(),
size=len(TINY_PNG),
)
signature = sign_payload(
{"temp": str(temp.id), "user": self.user.id}, services.UPLOAD_FILE_SALT
)
response = self.client.get(
f"/api/uploads/{temp.id}/file/", {"sig": signature}
)
self.assertEqual(response.status_code, 200)
self.assertIn(
f"max-age={services.TEMP_CACHE_SECONDS}", response["Cache-Control"]
)
self.assertNotIn("immutable", response["Cache-Control"])
def test_legacy_timestamp_signatures_are_accepted(self):
"""URLs minted before the stable scheme must not 500.
A TimestampSigner HMAC also passes a plain Signer's check, so the
embedded timestamp used to reach the JSON decoder and blow up.
"""
payload = {"item": self.item.id, "user": self.user.id, "action": "raw"}
legacy = signing.dumps(payload, salt=services.MEDIA_FILE_SALT)
self.assertEqual(
load_payload(legacy, services.MEDIA_FILE_SALT)["item"], self.item.id
)
response = self.client.get(
f"/api/files/J-{self.item.id}/raw/", {"sig": legacy}
)
self.assertEqual(response.status_code, 200)
def test_expired_and_malformed_signatures_return_none(self):
payload = {"item": self.item.id, "user": self.user.id, "action": "raw"}
with mock.patch.object(signing, "time") as clock:
clock.time.return_value = time.time() - 3 * 86400
expired = signing.dumps(payload, salt=services.MEDIA_FILE_SALT)
self.assertIsNone(load_payload(expired, services.MEDIA_FILE_SALT))
self.assertIsNone(load_payload("bogus", services.MEDIA_FILE_SALT))
self.assertIsNone(load_payload("a:b", services.MEDIA_FILE_SALT))
self.assertIsNone(load_payload("", services.MEDIA_FILE_SALT))
+401 -15
View File
@@ -6,17 +6,21 @@ import io
import json
import shutil
import tempfile
from datetime import timedelta
from pathlib import Path
from unittest import mock
from django.contrib.auth import get_user_model
from django.core.files.uploadedfile import SimpleUploadedFile
from django.test import Client, TestCase, override_settings
from django.utils import timezone
from PIL import Image
from rest_framework.authtoken.models import Token
from apps.library.models import MediaItem, TempUpload
from apps.library import upload_pipeline
from apps.library.models import MediaItem, TempUpload, UploadRun
User = get_user_model()
@@ -95,6 +99,7 @@ class TempUploadListTests(TestCase):
self.assertEqual(Client().get("/api/uploads/").status_code, 401)
@override_settings(UPLOAD_PIPELINE_AUTOSTART=False)
class StagedUploadWorkflowTests(TestCase):
@classmethod
def setUpClass(cls):
@@ -165,12 +170,18 @@ class StagedUploadWorkflowTests(TestCase):
self.assertEqual(len(body["resolved"]), 2)
self.assertEqual(body["errors"], [])
for temp in (first, second):
temp.refresh_from_db()
self.assertEqual(temp.status, TempUpload.STATUS_COMPLETED)
self.assertIsNotNone(temp.library_item_id)
self.assertEqual(temp.library_item.rating, "q")
self.assertEqual(temp.library_item.uploaded_by_id, self.uploader.id)
# Completed uploads are notifications now: the staged rows are gone
# and the live feed carries the filename -> J-ID mapping.
self.assertFalse(
TempUpload.objects.filter(pk__in=[first.id, second.id]).exists()
)
feed = UploadRun.objects.get(user=self.uploader).recent_completions
self.assertEqual(
{entry["filename"] for entry in feed}, {"one.png", "two.png"}
)
for item in MediaItem.objects.all():
self.assertEqual(item.rating, "q")
self.assertEqual(item.uploaded_by_id, self.uploader.id)
untouched.refresh_from_db()
self.assertEqual(untouched.status, TempUpload.STATUS_PENDING)
@@ -226,6 +237,25 @@ class StagedUploadWorkflowTests(TestCase):
self.assertEqual(response.status_code, 200)
return seed
def test_duplicate_upload_returns_a_completion_without_a_record(self):
client = self.api_client(self.uploader)
first = self.upload_via_api(client, "same.png")
self.assertEqual(first.status_code, 201)
# Index the first upload so the second one is byte-identical.
seed = TempUpload.objects.get(pk=first.json()["temp_id"])
self.resolve_bulk(client, [seed.id], "s")
response = self.upload_via_api(client, "same.png")
self.assertEqual(response.status_code, 201)
body = response.json()
self.assertEqual(body["status"], TempUpload.STATUS_COMPLETED)
self.assertEqual(body["resolution"], TempUpload.RESOLUTION_DUPLICATE)
self.assertTrue(body["library_j_id"].startswith("J-"))
# Duplicates never become board records; the feed announces them.
self.assertFalse(TempUpload.objects.filter(pk=body["temp_id"]).exists())
feed = UploadRun.objects.get(user=self.uploader).recent_completions
self.assertEqual(feed[-1]["filename"], "same.png")
def test_upload_defers_visual_similarity_to_its_phase(self):
client = self.api_client(self.uploader)
self.seed_library_item(client, "seed-defer")
@@ -257,6 +287,37 @@ class StagedUploadWorkflowTests(TestCase):
self.assertEqual(body["visual_matches"], [])
self.assertEqual(body["status"], TempUpload.STATUS_PENDING)
def test_detail_resigns_stored_visual_match_urls(self):
"""Stored matches carry only the item reference; URLs are re-minted.
Embedding the signed URL meant it expired (or used an older signing
scheme) and the modal showed alt text instead of thumbnails.
"""
client = self.api_client(self.uploader)
self.seed_library_item(client, "seed-resign")
item = MediaItem.objects.get()
temp = self.make_temp(self.uploader, "resign")
TempUpload.objects.filter(pk=temp.pk).update(
visual_matches=[
{
"item_id": item.id,
"j_id": f"J-{item.id}",
"filename": "seed-resign.png",
"similarity": 96.5,
"thumbnail_url": "/api/files/J-x/thumbnail/?sig=stale",
},
{"item_id": 999999, "j_id": "J-999999", "filename": "gone.png"},
]
)
body = client.get(f"/api/uploads/{temp.id}/").json()
self.assertEqual(len(body["visual_matches"]), 1)
match = body["visual_matches"][0]
self.assertEqual(match["j_id"], f"J-{item.id}")
self.assertEqual(match["similarity"], 96.5)
self.assertNotIn("stale", match["thumbnail_url"])
self.assertIn("/thumbnail/", match["thumbnail_url"])
self.assertIn(f"v={item.md5}", match["thumbnail_url"])
def test_visual_match_phase_rejects_completed_uploads(self):
client = self.api_client(self.uploader)
temp = self.make_temp(
@@ -307,14 +368,23 @@ class StagedUploadWorkflowTests(TestCase):
body = response.json()
self.assertEqual(body["errors"], [])
self.assertEqual(len(body["updated"]), 2)
for temp, post_id in ((first, 900001), (second, 900002)):
temp.refresh_from_db()
self.assertEqual(temp.status, TempUpload.STATUS_COMPLETED)
self.assertEqual(temp.library_item_id is not None, True)
self.assertEqual(temp.e621_post_id, post_id)
self.assertEqual(temp.resolution, TempUpload.RESOLUTION_AUTO_MD5)
self.assertEqual(temp.library_item.e621_post_id, post_id)
self.assertEqual(MediaItem.objects.count(), 2)
for entry, post_id in zip(body["updated"], (900001, 900002)):
self.assertEqual(entry["resolution"], TempUpload.RESOLUTION_AUTO_MD5)
self.assertEqual(entry["e621_post_id"], post_id)
self.assertTrue(entry["library_j_id"].startswith("J-"))
# Indexed uploads no longer leave a board record; the completion feed
# carries them for the live page instead.
self.assertFalse(TempUpload.objects.exists())
self.assertEqual(
{item.e621_post_id for item in MediaItem.objects.all()},
{900001, 900002},
)
feed = UploadRun.objects.get(user=self.uploader).recent_completions
self.assertEqual(len(feed), 2)
self.assertEqual(
{entry["filename"] for entry in feed},
{"bulk-link-1.png", "bulk-link-2.png"},
)
class IqdbRecordingTests(TestCase):
@@ -400,3 +470,319 @@ class IqdbRecordingTests(TestCase):
self.client, f"/api/uploads/{temp.id}/iqdb/", {"results": "nope"}
)
self.assertEqual(response.status_code, 400)
@override_settings(UPLOAD_PIPELINE_AUTOSTART=False)
class UploadPipelineTests(TestCase):
"""The server-side MD5 -> visual -> IQDB queue and its board API."""
@classmethod
def setUpClass(cls):
super().setUpClass()
cls._tmp = tempfile.mkdtemp(prefix="j621-pipeline-")
cls._watched = Path(cls._tmp) / "library"
cls._watched.mkdir(parents=True, exist_ok=True)
cls._settings = override_settings(
MEDIA_ROOT=cls._tmp, WATCHED_FOLDER=str(cls._watched)
)
cls._settings.enable()
@classmethod
def tearDownClass(cls):
cls._settings.disable()
shutil.rmtree(cls._tmp, ignore_errors=True)
super().tearDownClass()
def setUp(self):
self.uploader = User.objects.create_user(
username="pipe-uploader", password="pipe-pass-123456"
)
self.uploader.role = "uploader"
self.uploader.save(update_fields=["role"])
self.other = User.objects.create_user(
username="pipe-other", password="pipe-pass-123456"
)
self.other.role = "uploader"
self.other.save(update_fields=["role"])
def api_client(self, user):
client = Client()
client.defaults["HTTP_AUTHORIZATION"] = (
f"Token {Token.objects.create(user=user).key}"
)
return client
def stage(self, user=None, label="file", payload=None, filename=None):
user = user or self.uploader
payload = payload or TINY_PNG
name = filename or f"{label}.png"
return TempUpload.objects.create(
user=user,
file=SimpleUploadedFile(name, payload, content_type="image/png"),
original_filename=name,
md5=hashlib.md5(payload + label.encode()).hexdigest(),
size=len(payload),
)
def test_md5_match_auto_imports_the_file(self):
temp = self.stage(label="match")
temp_id = temp.id
post = {
"id": 123456,
"rating": "s",
"file": {
"md5": temp.md5,
"url": "https://static1.e621.net/data/m.png",
},
"tags": {"general": ["canine"]},
}
with mock.patch.object(
upload_pipeline.e621,
"check_md5_batch",
return_value={temp.md5: post},
):
upload_pipeline.run_pipeline(self.uploader.id)
# The indexed upload leaves no board record; the feed reports it.
self.assertFalse(TempUpload.objects.filter(pk=temp_id).exists())
item = MediaItem.objects.get(e621_post_id=123456)
self.assertEqual(item.uploaded_by_id, self.uploader.id)
run = UploadRun.objects.get(user=self.uploader)
self.assertEqual(run.status, UploadRun.STATUS_IDLE)
self.assertEqual(run.matched, 1)
self.assertEqual(run.processed, 1)
self.assertEqual(len(run.recent_completions), 1)
entry = run.recent_completions[0]
self.assertEqual(entry["id"], str(temp_id))
self.assertEqual(entry["item_id"], item.id)
self.assertEqual(entry["filename"], "match.png")
self.assertEqual(entry["resolution"], TempUpload.RESOLUTION_AUTO_MD5)
def test_status_reports_the_completion_feed(self):
temp = self.stage(label="feed")
client = self.api_client(self.uploader)
post = {
"id": 654321,
"rating": "s",
"file": {"md5": temp.md5, "url": "https://static1.e621.net/data/f.png"},
}
with mock.patch.object(
upload_pipeline.e621,
"check_md5_batch",
return_value={temp.md5: post},
):
upload_pipeline.run_pipeline(self.uploader.id)
body = client.get("/api/uploads/status/").json()
completions = body["recent_completions"]
self.assertEqual(len(completions), 1)
self.assertEqual(completions[0]["filename"], "feed.png")
self.assertEqual(completions[0]["j_id"], f"J-{MediaItem.objects.get().id}")
self.assertIn("/thumbnail/", completions[0]["thumbnail_url"])
def test_unmatched_file_runs_every_phase(self):
temp = self.stage(label="nomatch")
raw_iqdb = [
{
"post_id": 777,
"score": 91.0,
"post": {
"id": 777,
"rating": "q",
"md5": "a" * 32,
"score": 5,
"fav_count": 2,
"image_width": 800,
"image_height": 600,
},
}
]
modern = [
{
"id": 777,
"rating": "q",
"fav_count": 4,
"score": {"total": 9},
"preview": {"url": "https://static1.e621.net/data/preview/x.jpg"},
"file": {"md5": "a" * 32, "width": 801, "height": 601},
"tags": {"general": ["canine", "solo"]},
}
]
with mock.patch.object(
upload_pipeline.e621, "check_md5_batch", return_value={}
), mock.patch.object(
upload_pipeline.e621, "iqdb_search", return_value=raw_iqdb
), mock.patch.object(
upload_pipeline.e621, "fetch_posts_by_ids", return_value=modern
):
upload_pipeline.run_pipeline(self.uploader.id)
temp.refresh_from_db()
self.assertIsNotNone(temp.e621_checked_at)
self.assertIsNotNone(temp.visual_checked_at)
self.assertEqual(temp.status, TempUpload.STATUS_VISUAL_MATCH)
self.assertEqual(len(temp.iqdb_data), 1)
candidate = temp.iqdb_data[0]
self.assertEqual(candidate["post_id"], 777)
self.assertEqual(candidate["score_total"], 9)
self.assertEqual(candidate["fav_count"], 4)
self.assertEqual(candidate["width"], 801)
self.assertEqual(
candidate["preview_url"], "https://static1.e621.net/data/preview/x.jpg"
)
self.assertEqual(candidate["tags_preview"], ["canine", "solo"])
run = UploadRun.objects.get(user=self.uploader)
self.assertEqual(run.status, UploadRun.STATUS_IDLE)
self.assertEqual(run.processed, 1)
def test_visual_match_flags_similar_library_items(self):
from apps.library.uploads import complete_temp_upload
seed = self.stage(label="seed")
complete_temp_upload(seed)
self.assertTrue(MediaItem.objects.exists())
temp = self.stage(label="similar")
with mock.patch.object(
upload_pipeline.e621, "check_md5_batch", return_value={}
), mock.patch.object(
upload_pipeline.e621, "iqdb_search", return_value=[]
):
upload_pipeline.run_pipeline(self.uploader.id)
temp.refresh_from_db()
self.assertGreaterEqual(len(temp.visual_matches or []), 1)
self.assertEqual(temp.status, TempUpload.STATUS_VISUAL_MATCH)
def test_iqdb_rate_limit_pauses_the_run(self):
temp = self.stage(label="throttled")
with mock.patch.object(
upload_pipeline.e621, "check_md5_batch", return_value={}
), mock.patch.object(
upload_pipeline.e621,
"iqdb_search",
side_effect=upload_pipeline.e621.E621RateLimited("Throttled"),
):
upload_pipeline.run_pipeline(self.uploader.id)
run = UploadRun.objects.get(user=self.uploader)
self.assertEqual(run.status, UploadRun.STATUS_PAUSED)
self.assertIn("e621", run.error)
temp.refresh_from_db()
self.assertEqual(temp.attempts, 1)
self.assertNotEqual(temp.pipeline_error, "")
self.assertIsNone(temp.iqdb_data)
self.assertIsNone(temp.claimed_at)
client = self.api_client(self.uploader)
response = jpost(client, f"/api/uploads/{temp.id}/retry/", {})
self.assertEqual(response.status_code, 200)
temp.refresh_from_db()
self.assertEqual(temp.attempts, 0)
self.assertEqual(temp.pipeline_error, "")
def test_failed_rows_stop_after_the_attempt_cap(self):
temp = self.stage(label="broken")
for _ in range(upload_pipeline.MAX_ATTEMPTS):
with mock.patch.object(
upload_pipeline.e621, "check_md5_batch", return_value={}
), mock.patch.object(
upload_pipeline.e621,
"iqdb_search",
side_effect=upload_pipeline.e621.E621Error("boom"),
):
upload_pipeline.run_pipeline(self.uploader.id)
temp.refresh_from_db()
self.assertEqual(temp.attempts, upload_pipeline.MAX_ATTEMPTS)
self.assertEqual(upload_pipeline.count_outstanding(self.uploader), 0)
run = UploadRun.objects.get(user=self.uploader)
self.assertGreaterEqual(run.failed, 1)
def test_status_process_and_compact_board_payload(self):
temp = self.stage(label="board")
client = self.api_client(self.uploader)
status = client.get("/api/uploads/status/").json()
self.assertEqual(status["status"], UploadRun.STATUS_IDLE)
self.assertTrue(status["active"])
self.assertEqual(status["outstanding"], 1)
self.assertEqual(status["waiting"]["md5"], 1)
self.assertEqual(client.post("/api/uploads/process/").status_code, 200)
rows = client.get("/api/uploads/").json()
self.assertEqual(len(rows), 1)
row = rows[0]
for key in (
"md5_checked",
"visual_checked",
"iqdb_checked",
"processing",
"similar_count",
"pipeline_error",
):
self.assertIn(key, row)
self.assertNotIn("e621_data", row)
self.assertNotIn("iqdb_data", row)
detail = client.get(f"/api/uploads/{temp.id}/").json()
self.assertIn("e621_data", detail)
self.assertIn("iqdb_data", detail)
def test_discard_bulk_removes_only_own_rows(self):
client = self.api_client(self.uploader)
first = self.stage(label="discard-a")
second = self.stage(label="discard-b")
theirs = self.stage(self.other, label="discard-theirs")
paths = [Path(first.file.path), Path(second.file.path)]
response = jpost(
client,
"/api/uploads/discard-bulk/",
{"temp_ids": [str(first.id), str(second.id), str(theirs.id)]},
)
self.assertEqual(response.status_code, 200)
body = response.json()
self.assertEqual(len(body["discarded"]), 2)
self.assertEqual(len(body["errors"]), 1)
self.assertEqual(body["errors"][0]["error"], "not found")
for path in paths:
self.assertFalse(path.exists())
self.assertFalse(
TempUpload.objects.filter(pk__in=[first.id, second.id]).exists()
)
self.assertTrue(TempUpload.objects.filter(pk=theirs.id).exists())
def test_retry_with_phase_rechecks_iqdb(self):
temp = self.stage(label="recheck")
TempUpload.objects.filter(pk=temp.pk).update(
iqdb_data=[{"post_id": 1}],
status=TempUpload.STATUS_VISUAL_MATCH,
)
client = self.api_client(self.uploader)
response = jpost(
client, f"/api/uploads/{temp.id}/retry/", {"phase": "iqdb"}
)
self.assertEqual(response.status_code, 200)
temp.refresh_from_db()
self.assertIsNone(temp.iqdb_data)
self.assertEqual(temp.status, TempUpload.STATUS_VISUAL_MATCH)
def test_stale_claims_are_released(self):
temp = self.stage(label="stale")
TempUpload.objects.filter(pk=temp.pk).update(
claimed_at=timezone.now() - upload_pipeline.STALE_CLAIM_AFTER
- timedelta(minutes=1)
)
UploadRun.objects.create(user=self.uploader, status=UploadRun.STATUS_RUNNING)
UploadRun.objects.filter(user=self.uploader).update(
updated_at=timezone.now() - upload_pipeline.STALE_RUN_AFTER
- timedelta(minutes=1)
)
released, paused = upload_pipeline.reap_stale_claims()
self.assertEqual(released, 1)
self.assertEqual(paused, 1)
temp.refresh_from_db()
self.assertIsNone(temp.claimed_at)
run = UploadRun.objects.get(user=self.uploader)
self.assertEqual(run.status, UploadRun.STATUS_PAUSED)
+648
View File
@@ -0,0 +1,648 @@
"""Background processing for staged uploads.
Each staged upload runs through three phases, in batches of 75 (the same
lookup size the original app used for its e621 MD5 cache command):
1. e621 MD5 lookup — one ``posts.json`` query per round; byte-identical
matches are imported straight into the library (``auto_md5``).
2. Local visual similarity — perceptual hashes are compared against the
whole library once per round.
3. e621 IQDB — reverse-image search for whatever is still unresolved.
The pipeline runs in a daemon thread started on demand (like the download and
match scans), so the browser can navigate away and the work keeps going.
Progress and pause/error state live in the ``UploadRun`` row; per-file state
lives on ``TempUpload``. Rows are claimed with ``SELECT ... FOR UPDATE SKIP
LOCKED`` so several gunicorn workers cannot process the same file, and stale
claims left by a recycled worker are reaped and picked up again.
"""
import logging
import threading
import time
from datetime import timedelta
from django.conf import settings
from django.contrib.auth import get_user_model
from django.db import connection, transaction
from django.db.models import F, Q
from django.utils import timezone
from . import e621, services
from .models import MediaItem, TempUpload, UploadRun
logger = logging.getLogger(__name__)
# Same batch size as the original app's e621 cache command.
MD5_BATCH_SIZE = 75
# Rows one worker round claims; also the MD5 query size.
CLAIM_SIZE = MD5_BATCH_SIZE
# A claimed row is assumed dead after this long and is queued again.
STALE_CLAIM_AFTER = timedelta(minutes=15)
# A run whose heartbeat stopped this long ago can be taken over.
STALE_RUN_AFTER = timedelta(minutes=15)
# Per-row failures before the pipeline stops retrying automatically.
MAX_ATTEMPTS = 3
# Round-level e621 retries before the run is paused.
ROUND_ATTEMPTS = 3
ROUND_RETRY_SECONDS = 20
# Completions kept in the live feed. The board only shows what happened while
# the page was open, so a bounded rolling window is plenty.
COMPLETION_FEED_LIMIT = 200
WORK_STATUSES = (TempUpload.STATUS_PENDING, TempUpload.STATUS_VISUAL_MATCH)
VIDEO_RE = r"\.(mp4|webm)$"
# A staged upload still needs work when any phase has not run yet. IQDB is
# skipped for videos, which never get iqdb_data, so they must not stay
# "outstanding" forever.
OUTSTANDING_Q = (
Q(e621_checked_at__isnull=True)
| Q(visual_checked_at__isnull=True)
| (Q(iqdb_data__isnull=True) & ~Q(original_filename__iregex=VIDEO_RE))
)
_running_lock = threading.Lock()
_running_users: set[int] = set()
class PipelinePaused(Exception):
"""A round-level failure that should pause the run instead of failing rows."""
def is_video(filename):
return bool(filename) and filename.lower().endswith((".mp4", ".webm"))
def is_finished(temp):
"""True when every phase this file needs has run."""
if temp.status == TempUpload.STATUS_COMPLETED:
return True
if temp.e621_checked_at is None or temp.visual_checked_at is None:
return False
return is_video(temp.original_filename) or temp.iqdb_data is not None
def outstanding_queryset(user):
return (
TempUpload.objects.filter(user=user, status__in=WORK_STATUSES)
.filter(OUTSTANDING_Q)
.filter(attempts__lt=MAX_ATTEMPTS)
)
def count_outstanding(user):
return outstanding_queryset(user).count()
def count_failed(user):
return (
TempUpload.objects.filter(user=user, status__in=WORK_STATUSES)
.filter(attempts__gte=MAX_ATTEMPTS)
.count()
)
def waiting_counts(user):
"""How many files are left per phase (phases overlap by design)."""
base = TempUpload.objects.filter(
user=user, status__in=WORK_STATUSES, attempts__lt=MAX_ATTEMPTS
)
return {
"md5": base.filter(e621_checked_at__isnull=True).count(),
"visual": base.filter(visual_checked_at__isnull=True).count(),
"iqdb": (
base.filter(iqdb_data__isnull=True)
.exclude(original_filename__iregex=VIDEO_RE)
.count()
),
}
def status_payload(user, request=None):
"""Cheap state for the shell/upload page to poll."""
run = UploadRun.objects.filter(user=user).first()
outstanding = count_outstanding(user)
failed = count_failed(user)
status = run.status if run is not None else UploadRun.STATUS_IDLE
total = run.total if run is not None else 0
processed = run.processed if run is not None else 0
# A paused run with nothing left to do is not "active" (the user may have
# resolved or discarded the failed rows); failed rows stay visible until
# they are retried or dismissed.
active = (
outstanding > 0 or failed > 0 or status == UploadRun.STATUS_RUNNING
)
return {
"status": status,
"active": active,
"phase": run.phase if run is not None else "",
"total": max(total, processed + failed),
"processed": processed,
"matched": run.matched if run is not None else 0,
"failed": max(failed, run.failed if run is not None else 0),
"error": run.error if run is not None else "",
"outstanding": outstanding,
"waiting": waiting_counts(user),
"recent_completions": completion_payload(run, user, request=request),
"updated_at": run.updated_at.isoformat() if run is not None else None,
}
def completion_payload(run, user, request=None):
"""The live completion feed: filename -> J-ID for freshly indexed uploads.
Completed ``TempUpload`` rows are deleted, so this is the only place the
board learns about them. It is a notification feed, not durable state:
bounded, never dismissed, and ignored by fresh page loads.
"""
entries = list(run.recent_completions or []) if run is not None else []
if not entries:
return []
ids = [entry.get("item_id") for entry in entries if entry.get("item_id")]
items = MediaItem.objects.in_bulk(ids)
out = []
for entry in reversed(entries): # newest first
item = items.get(entry.get("item_id"))
if item is None:
continue
out.append(
{
"id": entry.get("id"),
"filename": entry.get("filename"),
"j_id": f"J-{item.id}",
"resolution": entry.get("resolution", ""),
"post_id": entry.get("post_id"),
"thumbnail_url": services.signed_media_url(
item, user, "thumbnail", request=request
),
"at": entry.get("at"),
}
)
return out
def record_completion(temp, item, resolution=""):
"""Append one completion to the owner's feed; never fails an import."""
entry = {
"id": str(temp.pk),
"filename": temp.original_filename,
"item_id": item.pk,
"resolution": resolution or temp.resolution or "",
"post_id": temp.e621_post_id,
"at": timezone.now().isoformat(),
}
try:
with transaction.atomic():
run, _ = UploadRun.objects.select_for_update().get_or_create(
user_id=temp.user_id
)
feed = list(run.recent_completions or [])
feed.append(entry)
run.recent_completions = feed[-COMPLETION_FEED_LIMIT:]
run.save(update_fields=["recent_completions", "updated_at"])
except Exception: # noqa: BLE001 - a notification must not break an import
logger.exception("Could not record the completion of %s", temp.pk)
def reap_stale_claims():
"""Queue rows left claimed by a recycled worker and pause dead runs."""
cutoff = timezone.now() - STALE_CLAIM_AFTER
released = TempUpload.objects.filter(claimed_at__lt=cutoff).update(
claimed_at=None
)
paused = UploadRun.objects.filter(
status=UploadRun.STATUS_RUNNING, updated_at__lt=cutoff
).update(
status=UploadRun.STATUS_PAUSED,
phase="",
error="The worker stopped before finishing. Retry to resume.",
updated_at=timezone.now(),
)
if released or paused:
logger.info(
"Reaped %s stale upload claims and %s dead upload runs", released, paused
)
return released, paused
def start_pipeline(user):
"""Start the pipeline for one user in a daemon thread (idempotent)."""
if not getattr(settings, "UPLOAD_PIPELINE_AUTOSTART", True):
return False
if user is None or not getattr(user, "can_upload", False):
return False
user_id = int(user.pk)
with _running_lock:
if user_id in _running_users:
return False
_running_users.add(user_id)
reap_stale_claims()
thread = threading.Thread(target=_thread_entry, args=(user_id,), daemon=True)
thread.start()
return True
def _thread_entry(user_id):
try:
run_pipeline(user_id)
except Exception: # noqa: BLE001 - a thread must never crash the worker
logger.exception("Upload pipeline for user %s crashed", user_id)
finally:
with _running_lock:
_running_users.discard(user_id)
connection.close()
def run_pipeline(user_id):
user = get_user_model().objects.filter(pk=user_id).first()
if user is None or not user.can_upload:
return
now = timezone.now()
# Claim the run row so two gunicorn workers cannot own the same queue.
with transaction.atomic():
run, _ = UploadRun.objects.select_for_update().get_or_create(user=user)
if (
run.status == UploadRun.STATUS_RUNNING
and run.updated_at is not None
and run.updated_at > now - STALE_RUN_AFTER
):
# Another worker owns this run.
return
# Reset the counters when a new queue starts cleanly; otherwise keep
# accumulating so failed rows from an earlier pass stay visible.
live_failed = count_failed(user)
outstanding = count_outstanding(user)
if run.status == UploadRun.STATUS_IDLE and live_failed == 0:
run.total = outstanding
run.processed = 0
run.matched = 0
run.failed = 0
else:
run.total = max(
run.total or 0, run.processed + run.failed + outstanding
)
run.failed = max(run.failed, live_failed)
run.status = UploadRun.STATUS_RUNNING
run.phase = ""
run.error = ""
run.started_at = now
run.save()
try:
while True:
rows = claim_round(user)
if not rows:
break
process_round(run, user, rows)
except PipelinePaused as exc:
_save_run(run, status=UploadRun.STATUS_PAUSED, phase="", error=str(exc))
except Exception as exc: # noqa: BLE001 - surface crashes as a run error
logger.exception("Upload pipeline for user %s failed", user_id)
_save_run(
run,
status=UploadRun.STATUS_ERROR,
phase="",
error=f"The upload pipeline stopped: {exc}",
)
else:
_save_run(run, status=UploadRun.STATUS_IDLE, phase="")
def claim_round(user, size=CLAIM_SIZE):
"""Claim up to ``size`` outstanding rows for this worker."""
now = timezone.now()
with transaction.atomic():
rows = list(
TempUpload.objects.select_for_update(skip_locked=True)
.filter(user=user, status__in=WORK_STATUSES)
.filter(OUTSTANDING_Q)
.filter(claimed_at__isnull=True, attempts__lt=MAX_ATTEMPTS)
.order_by("created_at")[:size]
)
if rows:
TempUpload.objects.filter(pk__in=[row.pk for row in rows]).update(
claimed_at=now
)
for row in rows:
row.claimed_at = now
return rows
def process_round(run, user, rows):
"""Run every phase for one claimed round, then release/account the rows."""
ids = [row.pk for row in rows]
matched = 0
try:
matched += md5_phase(run, user, rows)
rows = refresh(ids)
visual_phase(run, user, rows)
rows = refresh(ids)
iqdb_phase(run, user, rows)
finally:
finalize_round(run, ids, matched)
def refresh(ids):
return list(TempUpload.objects.filter(pk__in=ids))
def _save_run(run, **fields):
for key, value in fields.items():
setattr(run, key, value)
run.save(
update_fields=[*fields.keys(), "updated_at"]
)
def md5_phase(run, user, rows):
"""One e621 MD5 batch query; matches are imported into the library."""
targets = [row for row in rows if row.e621_checked_at is None]
if not targets:
return 0
_save_run(
run,
phase=UploadRun.PHASE_MD5,
total=run.processed + run.failed + count_outstanding(user),
)
posts = _e621_round(
lambda: e621.check_md5_batch(user, [row.md5 for row in targets])
)
by_md5 = {}
for post in posts.values():
file_data = post.get("file") or {}
md5 = str(file_data.get("md5") or "").strip().lower()
if md5:
by_md5[md5] = post
matched = 0
now = timezone.now()
for row in targets:
post = by_md5.get(str(row.md5).strip().lower())
if post is None:
TempUpload.objects.filter(pk=row.pk).update(
e621_checked_at=now, pipeline_error="", updated_at=now
)
continue
try:
trimmed = services.trim_e621_post(post)
if trimmed is None or not trimmed.get("id"):
raise e621.E621Error("e621 returned an unexpected post payload.")
TempUpload.objects.filter(pk=row.pk).update(
e621_post_id=int(trimmed["id"]),
e621_data=trimmed,
resolution=TempUpload.RESOLUTION_AUTO_MD5,
e621_checked_at=now,
pipeline_error="",
updated_at=now,
)
row.refresh_from_db()
from .uploads import complete_temp_upload
complete_temp_upload(row)
matched += 1
except Exception as exc: # noqa: BLE001 - keep going for other files
logger.exception("Could not auto-import staged upload %s", row.pk)
record_failure(row, f"Could not finish the upload: {exc}")
return matched
def visual_phase(run, user, rows):
"""Compare each row's perceptual hashes against the library once."""
from .uploads import build_hash_index, match_hashes
targets = [
row
for row in rows
if row.visual_checked_at is None
and row.status in WORK_STATUSES
and row.file
]
if not targets:
return
_save_run(run, phase=UploadRun.PHASE_VISUAL)
index = build_hash_index()
now = timezone.now()
for row in targets:
try:
hashes = services.compute_visual_hashes(row.file.path)
if not hashes:
TempUpload.objects.filter(pk=row.pk).update(
visual_checked_at=now, pipeline_error="", updated_at=now
)
continue
matches = match_hashes(hashes, index, user=user)
update = {
"visual_matches": matches,
"visual_checked_at": now,
"pipeline_error": "",
"updated_at": now,
}
if matches and row.status == TempUpload.STATUS_PENDING:
update["status"] = TempUpload.STATUS_VISUAL_MATCH
TempUpload.objects.filter(pk=row.pk).update(**update)
except Exception as exc: # noqa: BLE001 - keep going for other files
logger.exception("Visual similarity failed for %s", row.pk)
record_failure(row, f"Visual similarity failed: {exc}")
def iqdb_phase(run, user, rows):
"""Reverse-image search every unresolved image, one e621 query each."""
targets = [
row
for row in rows
if row.iqdb_data is None
and row.status in WORK_STATUSES
and row.file
and not is_video(row.original_filename)
]
if not targets:
return
_save_run(run, phase=UploadRun.PHASE_IQDB)
heartbeat_at = time.monotonic()
for row in targets:
# Keep the run row fresh: a 75-file IQDB round takes minutes and must
# not look like a dead worker to another request.
if time.monotonic() - heartbeat_at > 30:
UploadRun.objects.filter(pk=run.pk).update(updated_at=timezone.now())
heartbeat_at = time.monotonic()
try:
raw = e621.iqdb_search(user, row.file.path)
results = normalize_iqdb_results(user, raw)
except e621.E621AuthError as exc:
record_failure(row, str(exc))
raise PipelinePaused(
"e621 rejected the credentials — fix them in Account and retry."
) from exc
except e621.E621RateLimited as exc:
record_failure(row, str(exc))
raise PipelinePaused(
"e621 is throttling IQDB right now; the queue will resume."
) from exc
except Exception as exc: # noqa: BLE001 - keep going for other files
logger.exception("IQDB search failed for %s", row.pk)
record_failure(row, f"IQDB search failed: {exc}")
continue
now = timezone.now()
update = {
"iqdb_data": results,
"pipeline_error": "",
"updated_at": now,
}
if results and row.status == TempUpload.STATUS_PENDING:
update["status"] = TempUpload.STATUS_VISUAL_MATCH
TempUpload.objects.filter(pk=row.pk).update(**update)
def record_failure(row, message):
"""Count one failed attempt against a row and queue it for a retry."""
TempUpload.objects.filter(pk=row.pk).update(
attempts=F("attempts") + 1,
pipeline_error=str(message)[:2000],
claimed_at=None,
updated_at=timezone.now(),
)
def finalize_round(run, ids, matched):
"""Account finished/failed rows and release the rest of the claims."""
rows = refresh(ids)
finished = {row.pk for row in rows if is_finished(row)}
failed = {row.pk for row in rows if row.attempts >= MAX_ATTEMPTS}
# Release every claim: finished rows must not keep looking "processing"
# to the board, and unfinished rows are re-queued for the next run.
release = [row.pk for row in rows if row.claimed_at is not None]
if release:
TempUpload.objects.filter(pk__in=release).update(claimed_at=None)
_save_run(
run,
# Completed rows are deleted as they are imported, so they cannot be
# seen in the refreshed rows; count the matches explicitly.
processed=run.processed + len(finished) + matched,
failed=run.failed + len(failed - finished),
matched=run.matched + matched,
)
def _e621_round(task):
"""Run a round-level e621 call, retrying through rate limits."""
last_error = None
for attempt in range(ROUND_ATTEMPTS):
try:
return task()
except e621.E621AuthError:
raise
except (e621.E621RateLimited, e621.E621Error) as exc:
last_error = exc
if attempt + 1 >= ROUND_ATTEMPTS:
break
delay = ROUND_RETRY_SECONDS * (attempt + 1)
logger.info("e621 round failed (%s); retrying in %ss", exc, delay)
time.sleep(delay)
raise PipelinePaused(
f"e621 is not answering right now ({last_error}); the queue will resume."
) from last_error
def flatten_tag_preview(tags, limit=8):
"""First few tag names from a modern post payload, like the SPA shows."""
if not isinstance(tags, dict):
return []
out = []
for values in tags.values():
if not isinstance(values, list):
continue
for tag in values:
if isinstance(tag, str) and tag not in out:
out.append(tag)
if len(out) >= limit:
return out
return out
def _legacy_iqdb_post(entry):
"""Unwrap the post payload embedded in a legacy IQDB match."""
post = entry.get("post")
if not isinstance(post, dict):
return {}
inner = post.get("posts")
return inner if isinstance(inner, dict) else post
def normalize_iqdb_results(user, raw_results):
"""Shape legacy IQDB matches like the SPA's E621IqdbCandidate entries.
The IQDB payload carries little post data, so candidates are enriched
with one batched ``id:`` lookup before they are stored.
"""
candidates = []
for entry in (raw_results or [])[:10]:
if not isinstance(entry, dict):
continue
post = _legacy_iqdb_post(entry)
post_id = entry.get("post_id")
if not isinstance(post_id, int):
post_id = post.get("id")
score = entry.get("score")
candidates.append(
{
"post_id": post_id if isinstance(post_id, int) else None,
"score": float(score) if isinstance(score, (int, float)) else None,
"preview_url": None,
"rating": (
post.get("rating") if isinstance(post.get("rating"), str) else None
),
"md5": post.get("md5") if isinstance(post.get("md5"), str) else None,
"score_total": (
post.get("score") if isinstance(post.get("score"), int) else None
),
"fav_count": (
post.get("fav_count")
if isinstance(post.get("fav_count"), int)
else None
),
"width": (
post.get("image_width")
if isinstance(post.get("image_width"), int)
else None
),
"height": (
post.get("image_height")
if isinstance(post.get("image_height"), int)
else None
),
"tags_preview": [],
}
)
ids = [entry["post_id"] for entry in candidates if entry["post_id"]]
if not ids:
return services.sanitize_iqdb_results(candidates)
try:
posts = e621.fetch_posts_by_ids(user, ids)
except e621.E621Error as exc:
# Candidates without enrichment still show up; keep them.
logger.info("Could not enrich IQDB candidates: %s", exc)
posts = []
by_id = {post.get("id"): post for post in posts}
for entry in candidates:
post = by_id.get(entry["post_id"])
if not isinstance(post, dict):
continue
file_data = post.get("file") or {}
preview = post.get("preview") or {}
score = post.get("score") or {}
entry["preview_url"] = preview.get("url") or entry["preview_url"]
entry["rating"] = post.get("rating") or entry["rating"]
entry["md5"] = file_data.get("md5") or entry["md5"]
if isinstance(score, dict):
entry["score_total"] = score.get("total")
entry["fav_count"] = post.get("fav_count")
entry["width"] = file_data.get("width")
entry["height"] = file_data.get("height")
entry["tags_preview"] = flatten_tag_preview(post.get("tags"))
return services.sanitize_iqdb_results(candidates)
+237 -38
View File
@@ -16,7 +16,6 @@ from urllib.parse import urlparse
from django.conf import settings
from django.contrib.auth import get_user_model
from django.core import signing
from django.core.exceptions import ValidationError
from django.http import Http404
from django.utils import timezone
@@ -29,32 +28,39 @@ from rest_framework.response import Response
from . import services
from .models import MediaItem, TempUpload
from .permissions import CanUpload
from .serializers import TempUploadSerializer
from .tools import HASH_FIELDS, hashes_similarity
from .serializers import TempUploadListSerializer, TempUploadSerializer
from .signing_urls import load_payload
from .tools import HASH_FIELDS, hashed_items, hashes_similarity
logger = logging.getLogger(__name__)
def find_library_matches(path, limit=10, user=None, request=None):
"""Library items visually similar to a staged file."""
hashes = services.compute_visual_hashes(path)
if not hashes:
return []
def build_hash_index():
"""Library hash mappings for similarity scans, loaded once per batch.
Only items that actually carry perceptual hashes are included; the old
per-file scan walked every row (including videos and unchecked items).
"""
algorithms = list(HASH_FIELDS)
return [
(item, {field: getattr(item, field, "") for field in algorithms})
for item in hashed_items(algorithms)
]
def match_hashes(hashes, index, limit=10, user=None, request=None):
"""Library items whose perceptual hashes are close to ``hashes``."""
algorithms = list(HASH_FIELDS)
threshold = settings.VISUAL_MATCH_THRESHOLD
matches = []
for item in MediaItem.objects.prefetch_related("locations"):
similarity = hashes_similarity(
hashes,
{field: getattr(item, field, "") for field in algorithms},
algorithms,
threshold,
)
for item, item_hashes in index:
similarity = hashes_similarity(hashes, item_hashes, algorithms, threshold)
if similarity is None:
continue
location = item.locations.first()
matches.append(
{
"item_id": item.id,
"j_id": f"J-{item.id}",
"filename": Path(location.rel_path).name if location else item.md5,
"similarity": round(similarity * 100, 1),
@@ -67,6 +73,20 @@ def find_library_matches(path, limit=10, user=None, request=None):
return matches[:limit]
def find_library_matches(path, limit=10, user=None, request=None, index=None):
"""Library items visually similar to a staged file.
Pass a prebuilt ``index`` (see build_hash_index) to reuse it across a
whole batch instead of rescanning the library per file.
"""
hashes = services.compute_visual_hashes(path)
if not hashes:
return []
if index is None:
index = build_hash_index()
return match_hashes(hashes, index, limit=limit, user=user, request=request)
def complete_temp_upload(temp, download_url=None):
"""Index the upload into the library.
@@ -108,6 +128,12 @@ def complete_temp_upload(temp, download_url=None):
services.ensure_visual_hashes(item)
temp.file.delete(save=False)
# Warm the preview while the import is still off the request path.
try:
services.ensure_thumbnail(item)
except Exception: # noqa: BLE001 - a preview must not fail the import
logger.exception("Could not warm the thumbnail for J-%s", item.id)
temp.library_item = item
temp.status = TempUpload.STATUS_COMPLETED
@@ -139,6 +165,13 @@ def complete_temp_upload(temp, download_url=None):
item.save(update_fields=update_fields + ["updated_at"])
temp.save()
from .upload_pipeline import record_completion
record_completion(temp, item)
# The board learns about completions from the live feed, so the record is
# deleted as soon as the file is indexed: nothing left to dismiss. Delete
# through the queryset so callers keep ``temp.pk`` for their response.
TempUpload.objects.filter(pk=temp.pk).delete()
return item
@@ -156,12 +189,42 @@ class TempUploadViewSet(
pagination_class = None
http_method_names = ["get", "post", "delete", "head", "options"]
def get_serializer_class(self):
# The board polls the list, so its payload stays small; the metadata
# modal fetches the full row from the detail endpoint.
if self.action == "list":
return TempUploadListSerializer
return TempUploadSerializer
def get_queryset(self):
queryset = TempUpload.objects.select_related("library_item")
user = self.request.user
if not user.is_app_staff:
queryset = queryset.filter(user=user)
return queryset
# Staged uploads are private: everyone, staff included, only sees
# their own board. (The file action still lets staff read bytes by id
# for support purposes.)
return TempUpload.objects.select_related("library_item").filter(
user=self.request.user
)
def _completed_payload(self, temp, item):
"""Synthetic row for an upload that is indexed immediately.
Duplicates and resolved uploads never leave a board record; the SPA
turns this response (or the live completion feed) into a "J-x
uploaded" card that lives only in the page session.
"""
return {
"temp_id": str(temp.pk),
"original_filename": temp.original_filename,
"md5": temp.md5,
"size": temp.size,
"status": TempUpload.STATUS_COMPLETED,
"resolution": temp.resolution,
"e621_post_id": temp.e621_post_id,
"library_j_id": f"J-{item.id}",
"file_url": None,
"preview_url": services.signed_media_url(
item, self.request.user, "thumbnail", request=self.request
),
}
def create(self, request):
upload = request.FILES.get("file")
@@ -190,10 +253,21 @@ class TempUploadViewSet(
temp.resolution = TempUpload.RESOLUTION_DUPLICATE
temp.library_item = existing
temp.file.delete(save=False)
# Visual similarity is deliberately a separate phase (the
# /visual-match action) so a large batch uploads at full speed and
# the board runs MD5 -> visual -> IQDB over the whole batch.
temp.save()
from .upload_pipeline import record_completion
record_completion(temp, existing)
payload = self._completed_payload(temp, existing)
TempUpload.objects.filter(pk=temp.pk).delete()
return Response(payload, status=status.HTTP_201_CREATED)
# Visual similarity and IQDB run in the background pipeline so a large
# batch uploads at full speed and the work survives the browser.
temp.save()
# Kick the server-side pipeline; staging no longer waits on e621 and
# the work continues even if the browser navigates away.
from .upload_pipeline import start_pipeline
start_pipeline(request.user)
return Response(
self.get_serializer(temp).data, status=status.HTTP_201_CREATED
)
@@ -221,21 +295,148 @@ class TempUploadViewSet(
temp.save(update_fields=["visual_matches", "status", "updated_at"])
return Response(self.get_serializer(temp).data)
@action(detail=True, methods=["get", "head"], permission_classes=[AllowAny])
@action(detail=False, methods=["get"])
def status(self, request):
"""Cheap pipeline state for the shell indicator and the upload page."""
from .upload_pipeline import status_payload
return Response(status_payload(request.user, request=request))
@action(detail=False, methods=["post"])
def process(self, request):
"""Start (or resume) the pipeline for the caller's staged uploads.
Idempotent: the client calls this after staging files, on page load
and when a paused run should be retried.
"""
from .upload_pipeline import start_pipeline, status_payload
start_pipeline(request.user)
return Response(status_payload(request.user, request=request))
@action(detail=True, methods=["post"])
def retry(self, request, pk=None):
"""Queue one staged upload for another pipeline pass."""
from .upload_pipeline import start_pipeline, status_payload
temp = self.get_object()
if temp.status == TempUpload.STATUS_COMPLETED:
return Response(
{"detail": "This upload is already in the library."},
status=status.HTTP_400_BAD_REQUEST,
)
if not temp.file:
return Response(
{"detail": "The staged file is missing."},
status=status.HTTP_400_BAD_REQUEST,
)
update = {
"claimed_at": None,
"attempts": 0,
"pipeline_error": "",
"updated_at": timezone.now(),
}
# An explicit phase re-runs that one check even if it already ran.
phase = str(request.data.get("phase") or "").strip()
if phase == "md5":
update["e621_checked_at"] = None
elif phase == "visual":
update["visual_matches"] = None
update["visual_checked_at"] = None
elif phase == "iqdb":
update["iqdb_data"] = None
if temp.status == TempUpload.STATUS_ERROR:
# A failed import has to go through the MD5 phase again so the
# completion is retried; other errors only re-run missing phases.
update["status"] = TempUpload.STATUS_PENDING
update["e621_checked_at"] = None
elif temp.status not in (
TempUpload.STATUS_PENDING,
TempUpload.STATUS_VISUAL_MATCH,
):
update["status"] = TempUpload.STATUS_PENDING
TempUpload.objects.filter(pk=temp.pk).update(**update)
start_pipeline(request.user)
return Response(status_payload(request.user, request=request))
@action(detail=False, methods=["post"], url_path="retry-all")
def retry_all(self, request):
"""Queue every retryable staged upload for another pipeline pass."""
from .upload_pipeline import start_pipeline, status_payload
now = timezone.now()
retryable = self.get_queryset().filter(
status__in=[
TempUpload.STATUS_PENDING,
TempUpload.STATUS_VISUAL_MATCH,
TempUpload.STATUS_ERROR,
]
)
retryable.exclude(file="").update(
claimed_at=None,
attempts=0,
pipeline_error="",
updated_at=now,
)
retryable.exclude(file="").filter(status=TempUpload.STATUS_ERROR).update(
status=TempUpload.STATUS_PENDING,
e621_checked_at=None,
)
start_pipeline(request.user)
return Response(status_payload(request.user, request=request))
@action(detail=False, methods=["post"], url_path="discard-bulk")
def discard_bulk(self, request):
"""Discard many staged uploads in one request (the board's "all")."""
ids = request.data.get("temp_ids")
if not isinstance(ids, list) or not ids:
return Response(
{"detail": "temp_ids must be a non-empty list."},
status=status.HTTP_400_BAD_REQUEST,
)
if len(ids) > 1000:
return Response(
{"detail": "Too many ids in one request (max 1000)."},
status=status.HTTP_400_BAD_REQUEST,
)
values = list(dict.fromkeys(str(value) for value in ids))
try:
queryset = self.get_queryset().filter(pk__in=values)
except (ValidationError, ValueError):
return Response(
{"detail": "One or more temp_ids are not valid upload ids."},
status=status.HTTP_400_BAD_REQUEST,
)
by_id = {str(temp.pk): temp for temp in queryset}
discarded: list[str] = []
errors: list[dict[str, str]] = []
for value in values:
temp = by_id.get(value)
if temp is None:
errors.append({"temp_id": value, "error": "not found"})
continue
try:
self.perform_destroy(temp)
discarded.append(value)
except Exception as exc: # noqa: BLE001 - report per-file failures
logger.exception("Could not discard staged upload %s", value)
errors.append({"temp_id": value, "error": str(exc)})
return Response({"discarded": discarded, "errors": errors})
@action(
detail=True,
methods=["get", "head"],
permission_classes=[AllowAny],
throttle_classes=[],
)
def file(self, request, pk=None):
"""Serve the staged file; accepts a signed URL for media tags."""
user = request.user if request.user.is_authenticated else None
if user is None:
signature = request.query_params.get("sig")
if signature:
try:
payload = signing.loads(
signature,
salt=services.UPLOAD_FILE_SALT,
max_age=86400,
)
except signing.BadSignature:
payload = None
payload = load_payload(signature, services.UPLOAD_FILE_SALT)
if payload and str(payload.get("temp")) == str(pk):
user = (
get_user_model()
@@ -336,7 +537,7 @@ class TempUploadViewSet(
temp.save()
try:
complete_temp_upload(temp, download_url=download_url)
item = complete_temp_upload(temp, download_url=download_url)
except Exception as exc: # noqa: BLE001 - report completion failures
logger.exception("Could not complete staged upload %s", temp.id)
temp.status = TempUpload.STATUS_ERROR
@@ -345,8 +546,7 @@ class TempUploadViewSet(
{"detail": f"Could not finish the upload: {exc}"},
status=status.HTTP_400_BAD_REQUEST,
)
temp.refresh_from_db()
return Response(self.get_serializer(temp).data)
return Response(self._completed_payload(temp, item))
@action(detail=False, methods=["post"], url_path="link-bulk")
def link_bulk(self, request):
@@ -418,15 +618,14 @@ class TempUploadViewSet(
candidate_url = ""
temp.save()
try:
complete_temp_upload(temp, download_url=candidate_url or None)
item = complete_temp_upload(temp, download_url=candidate_url or None)
except Exception as exc: # noqa: BLE001 - report per-file failures
logger.exception("Could not complete staged upload %s", temp.id)
temp.status = TempUpload.STATUS_ERROR
temp.save(update_fields=["status", "updated_at"])
errors.append({"temp_id": temp_id, "error": str(exc)})
continue
temp.refresh_from_db()
updated.append(self.get_serializer(temp).data)
updated.append(self._completed_payload(temp, item))
return Response({"updated": updated, "errors": errors})
+26 -12
View File
@@ -5,7 +5,6 @@ from pathlib import Path
from urllib.parse import urlparse
from django.conf import settings
from django.core import signing
from django.db.models import Min, Q
from django.http import Http404, StreamingHttpResponse
from django.shortcuts import get_object_or_404
@@ -31,6 +30,7 @@ from .serializers import (
MatchTaskSerializer,
MediaItemSerializer,
)
from .signing_urls import load_payload
LIST_ORDERINGS = {"name", "-name", "size", "-size", "created_at", "-created_at"}
MD5_RE = re.compile(r"[0-9a-fA-F]{32}")
@@ -132,11 +132,8 @@ class MediaItemViewSet(
signature = request.query_params.get("sig")
if not signature:
return None
try:
payload = signing.loads(
signature, salt=services.MEDIA_FILE_SALT, max_age=86400
)
except signing.BadSignature:
payload = load_payload(signature, services.MEDIA_FILE_SALT)
if payload is None:
return None
if payload.get("action") != action_name:
return None
@@ -148,7 +145,7 @@ class MediaItemViewSet(
return item
return self.get_object()
@action(detail=True, methods=["get"])
@action(detail=True, methods=["get"], throttle_classes=[])
def raw(self, request, pk=None):
item = self._media_object(request, "raw")
location = item.locations.first()
@@ -158,10 +155,14 @@ class MediaItemViewSet(
status=status.HTTP_404_NOT_FOUND,
)
return services.serve_file(
request, location.path, download=request.query_params.get("download") == "1"
request,
location.path,
download=request.query_params.get("download") == "1",
max_age=services.MEDIA_CACHE_SECONDS,
immutable=True,
)
@action(detail=True, methods=["get"])
@action(detail=True, methods=["get"], throttle_classes=[])
def thumbnail(self, request, pk=None):
item = self._media_object(request, "thumbnail")
location = item.locations.first()
@@ -173,13 +174,26 @@ class MediaItemViewSet(
path = Path(location.path)
if path.suffix.lower() in services.VIDEO_EXTENSIONS:
thumbnail = services.generate_video_thumbnail(item.md5, path)
if thumbnail is None:
else:
thumbnail = services.generate_image_thumbnail(item.md5, path)
if thumbnail is not None:
return services.serve_file(
request,
thumbnail,
max_age=services.MEDIA_CACHE_SECONDS,
immutable=True,
)
if path.suffix.lower() in services.VIDEO_EXTENSIONS:
return Response(
{"detail": "Thumbnail unavailable."},
status=status.HTTP_404_NOT_FOUND,
)
return services.serve_file(request, thumbnail)
return services.serve_file(request, path)
return services.serve_file(
request,
path,
max_age=services.MEDIA_CACHE_SECONDS,
immutable=True,
)
@action(detail=False, methods=["post"], permission_classes=[AllowAny])
def lookup(self, request):
+26 -3
View File
@@ -234,6 +234,12 @@ GUEST_BLACKLIST_TTL = int(os.getenv("GUEST_BLACKLIST_TTL", "3600"))
# Similarity threshold for flagging staged uploads that match library items.
VISUAL_MATCH_THRESHOLD = float(os.getenv("VISUAL_MATCH_THRESHOLD", "0.9"))
# Start the staged-upload pipeline when a file is staged (daemon thread in the
# worker). Tests turn this off and drive the pipeline synchronously.
UPLOAD_PIPELINE_AUTOSTART = os.getenv(
"UPLOAD_PIPELINE_AUTOSTART", "true"
).strip().lower() not in {"0", "false", "no", "off"}
# Ephemeral similarity-check uploads are deleted after this many minutes
# (and always on startup).
SIMILARITY_TTL_MINUTES = int(os.getenv("SIMILARITY_TTL_MINUTES", "30"))
@@ -249,6 +255,16 @@ CACHES = {
# Django REST Framework
# Private / tailnet-only deployments can drop the general anon+user limits
# entirely (THROTTLE_ENABLED=false). The scoped guards below (login, register,
# e621 proxy) and the media endpoints' own protections stay active either way.
THROTTLE_ENABLED = os.getenv("THROTTLE_ENABLED", "true").strip().lower() not in {
"0",
"false",
"no",
"off",
}
REST_FRAMEWORK = {
"DEFAULT_AUTHENTICATION_CLASSES": [
"rest_framework.authentication.TokenAuthentication",
@@ -263,11 +279,18 @@ REST_FRAMEWORK = {
],
"DEFAULT_PAGINATION_CLASS": "config.pagination.StandardPagination",
"PAGE_SIZE": 48,
# Per-IP/per-user rate limits (counted in the shared Redis cache).
"DEFAULT_THROTTLE_CLASSES": [
# Per-IP/per-user rate limits (counted in the shared Redis cache). Signed
# media URLs are deliberately excluded at the view level: <img>/<video>
# tags fetch them without an Authorization header, so a library page would
# otherwise burn the anonymous bucket and start returning JSON 429s.
"DEFAULT_THROTTLE_CLASSES": (
[
"rest_framework.throttling.AnonRateThrottle",
"rest_framework.throttling.UserRateThrottle",
],
]
if THROTTLE_ENABLED
else []
),
"DEFAULT_THROTTLE_RATES": {
# Generous enough for the shell polling (status every 5s, stats every 2s).
"anon": os.getenv("THROTTLE_ANON", "120/min"),
+4
View File
@@ -64,6 +64,10 @@ DB_ROOT_PASSWORD=j621root
# THROTTLE_LOGIN=5/min
# THROTTLE_REGISTER=20/hour
# THROTTLE_E621_PROXY=60/hour
# Tailnet-only / private deployments can drop the general limits entirely.
# Signed media URLs (<img>/<video>) and the login/register/proxy guards are
# exempt from this switch either way.
# THROTTLE_ENABLED=false
# e621 media hosts the backend may fetch from (downloads, proxies)
# E621_MEDIA_HOSTS=static1.e621.net,static2.e621.net,static3.e621.net
+12 -5
View File
@@ -145,9 +145,12 @@ Build the desktop installers without publishing anything:
```
Hand the files out or attach them to a Gitea release manually — the script
prints sizes and SHA-256 sums for the release notes.
prints sizes and SHA-256 sums for the release notes. The release assets are
also the desktop update feed; the app resolves the newest `desktop-v*`
release on Gitea at check time (see `desktop/README.md`).
Push the desktop builds and their update metadata to the frontend's feed:
The frontend's `/desktop/` feed is optional now — kept for manual downloads
and for installs older than 0.1.2. To publish it:
```bash
./push_desktop.sh # build Linux packages + copy the feed to jakerasp
@@ -160,9 +163,13 @@ The remote copy defaults to `jakerasp:/home/jake/servers/J621`, or
`$J621_DESKTOP_FEED_HOST` when set. Artifacts land in `deploy/data/desktop/`,
which the frontend nginx mounts read-only and serves at `/desktop/`. The
remote copy uses rsync when both ends have it, tar over ssh when the server
does not. The desktop app's "Check for updates…" menu item reads
`latest-linux.yml` / `latest.yml` from there (see `desktop/README.md`).
Backend-only composes have no frontend, so no feed.
does not. Backend-only composes have no frontend, so no website feed.
The manual **CD** workflow (Actions tab) builds the desktop packages on the
runner and attaches them plus the update metadata to the Gitea release
`desktop-v<version>`; that is what the desktop updater reads. The website feed
is not touched by CI (runtime state on the deploy host, no SSH key there); use
`push_desktop.sh` when it needs refreshing for old installs.
## Scheduled jobs
+15 -2
View File
@@ -22,6 +22,8 @@ case "$TARGET" in
;;
esac
VERSION="$(node -p "require('./desktop/package.json').version")"
if [ ! -d desktop/node_modules ]; then
echo "==> Installing desktop dependencies ..."
npm --prefix desktop ci --no-audit --no-fund
@@ -32,14 +34,25 @@ if [ "$TARGET" != "linux" ] && ! command -v wine >/dev/null 2>&1; then
exit 1
fi
# release/ should only ever hold the current version: old installers are
# rebuilt from git when needed, and the unpacked trees are regenerated.
shopt -s nullglob
stale=(desktop/release/*.deb desktop/release/*.pkg.tar.zst desktop/release/"J621 Setup "*.exe
desktop/release/*.blockmap desktop/release/latest*.yml)
for file in "${stale[@]}"; do
case "$(basename "$file")" in
*"$VERSION"*) ;;
*) echo " removing $(basename "$file")"; rm -f "$file" ;;
esac
done
rm -rf desktop/release/linux-unpacked desktop/release/win-unpacked
case "$TARGET" in
linux) npm --prefix desktop run dist:linux ;;
win) npm --prefix desktop run dist:win ;;
all) npm --prefix desktop run dist:all ;;
esac
VERSION="$(node -p "require('./desktop/package.json').version")"
case "$TARGET" in
linux) ARTIFACTS=(-name "*${VERSION}*.deb" -o -name "*${VERSION}*.pkg.tar.zst") ;;
win) ARTIFACTS=(-name "*${VERSION}*.exe") ;;
+8 -2
View File
@@ -10,10 +10,16 @@ REGISTRY="gitea.rainbow-herring.ts.net/jakebreath/j621-backend"
REGISTRY_HOST="$(printf '%s' "$REGISTRY" | cut -d/ -f1)"
SHA="${1:-$(git rev-parse --short HEAD)}"
BUILDER=multiarch
PLATFORMS="linux/amd64,linux/arm64"
PLATFORMS="${PLATFORMS:-linux/amd64,linux/arm64}"
echo "==> Logging in to $REGISTRY_HOST ..."
docker login "$REGISTRY_HOST"
if [ -n "${REGISTRY_USER:-}" ] && [ -n "${REGISTRY_TOKEN:-}" ]; then
# Non-interactive login for CI (workflow passes GITHUB_TOKEN).
printf '%s' "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" \
-u "$REGISTRY_USER" --password-stdin
else
docker login "$REGISTRY_HOST"
fi
if ! docker buildx inspect "$BUILDER" >/dev/null 2>&1; then
echo "==> Creating buildx builder '$BUILDER' ..."
+17 -10
View File
@@ -1,11 +1,12 @@
#!/bin/bash
# Build the J621 desktop packages and publish them to the update feed.
# Build the J621 desktop packages, optionally publish them to the website
# feed.
#
# Locally the feed is deploy/data/desktop, which the frontend nginx mounts
# read-only and serves at /desktop/. With --host the same directory is also
# copied to a remote deploy checkout (rsync, or tar over ssh when the server
# has no rsync). electron-updater reads latest-linux.yml / latest.yml from
# there; the feed URL comes from desktop/electron-builder.yml.
# Desktop updates no longer depend on this: the app resolves the newest
# `desktop-v*` release on Gitea at check time (see desktop/README.md). This
# script builds the packages and can copy them to deploy/data/desktop, which
# the frontend nginx mounts read-only and serves at /desktop/ for manual
# downloads and for pre-0.1.2 installs.
#
# Usage: ./push_desktop.sh [--win] [--no-build] [--local] [--host user@server:/path]
# --win also cross-build the Windows NSIS installer (needs wine)
@@ -92,6 +93,16 @@ if [ -n "$REMOTE" ]; then
echo
echo "==> Copying the feed to $REMOTE_TARGET:$REMOTE_FEED ..."
if ! ssh "$REMOTE_TARGET" "mkdir -p '$REMOTE_FEED'"; then
echo "ssh to $REMOTE_TARGET failed; nothing was copied." >&2
exit 1
fi
# The remote feed keeps only the current version, like the local one.
if ! ssh "$REMOTE_TARGET" \
"find '$REMOTE_FEED' -maxdepth 1 -type f \\( -name '*.deb' -o -name '*.pkg.tar.zst' -o -name '*.exe' -o -name '*.exe.blockmap' \\) ! -name '*$VERSION*' -exec rm -f {} +"; then
echo "==> Could not prune old installers on $REMOTE_TARGET (continuing)." >&2
fi
transferred=0
if command -v rsync >/dev/null 2>&1; then
if rsync -a --info=progress2 "$FEED"/ "$REMOTE_TARGET:$REMOTE_FEED/"; then
@@ -101,10 +112,6 @@ if [ -n "$REMOTE" ]; then
fi
fi
if [ "$transferred" -eq 0 ]; then
if ! ssh "$REMOTE_TARGET" "mkdir -p '$REMOTE_FEED'"; then
echo "ssh to $REMOTE_TARGET failed; nothing was copied." >&2
exit 1
fi
if ! tar -C "$FEED" -cf - . |
ssh "$REMOTE_TARGET" "tar -C '$REMOTE_FEED' -xf -"; then
echo "Copy to $REMOTE_TARGET:$REMOTE_FEED failed." >&2
+8 -2
View File
@@ -10,10 +10,16 @@ REGISTRY="gitea.rainbow-herring.ts.net/jakebreath/j621-frontend"
REGISTRY_HOST="$(printf '%s' "$REGISTRY" | cut -d/ -f1)"
SHA="${1:-$(git rev-parse --short HEAD)}"
BUILDER=multiarch
PLATFORMS="linux/amd64,linux/arm64"
PLATFORMS="${PLATFORMS:-linux/amd64,linux/arm64}"
echo "==> Logging in to $REGISTRY_HOST ..."
docker login "$REGISTRY_HOST"
if [ -n "${REGISTRY_USER:-}" ] && [ -n "${REGISTRY_TOKEN:-}" ]; then
# Non-interactive login for CI (workflow passes GITHUB_TOKEN).
printf '%s' "$REGISTRY_TOKEN" | docker login "$REGISTRY_HOST" \
-u "$REGISTRY_USER" --password-stdin
else
docker login "$REGISTRY_HOST"
fi
if ! docker buildx inspect "$BUILDER" >/dev/null 2>&1; then
echo "==> Creating buildx builder '$BUILDER' ..."
+19 -8
View File
@@ -49,7 +49,8 @@ npm run dist:all
`deploy/build_desktop.sh` wraps the same commands, installs dependencies on
first run and prints sizes plus SHA-256 sums for release notes. Nothing is
published by it; `deploy/push_desktop.sh` is the one that feeds auto-updates.
published by it; the CD workflow attaches the artifacts to the Gitea release,
which is also the update feed.
The Arch package can be installed and removed with pacman:
@@ -68,19 +69,29 @@ will warn, and it has not been smoke-tested on real Windows.
The app checks only when asked (**J621 → Check for updates…** in the menu):
Linux packages install through pacman/dpkg, which needs administrator rights,
and the Windows build is unsigned, so nothing installs silently. The check
reads `latest-linux.yml` / `latest.yml` from the feed configured in
`electron-builder.yml` (`publish.url`, baked into `app-update.yml`); set
`J621_UPDATE_URL` to point a build at another feed (the smoke test uses this).
and the Windows build is unsigned, so nothing installs silently.
The check resolves the feed itself: it asks the Gitea API for the newest
non-draft `desktop-v*` release (`J621_UPDATE_REPO`, default
`https://gitea.rainbow-herring.ts.net/jakebreath/j621` — lowercase on purpose,
the API path is case-sensitive), picks the `latest-linux.yml` / `latest.yml`
asset for the platform and uses that release as an electron-updater generic
feed. The baked `publish.url` in `electron-builder.yml` is metadata only.
`J621_UPDATE_URL` overrides the whole lookup (the smoke test uses this).
Publishing a release:
1. Bump `version` in `desktop/package.json` — that is what the updater compares.
2. `./deploy/push_desktop.sh --win` builds deb/pacman/NSIS and copies the
artifacts plus both channel files into `deploy/data/desktop/`, which the
frontend nginx serves read-only at `/desktop/`.
2. Run the manual CD workflow, which builds the packages and attaches them
plus both channel files to the Gitea release `desktop-v<version>`.
`./deploy/push_desktop.sh --win` does the same build locally (and can also
copy the files to the website feed, which is optional now).
3. Existing installs find the new version on their next manual check.
Note for the 0.1.1 → 0.1.2 step: 0.1.1 only knows the old `/desktop/` feed, so
publish 0.1.2 there once (`./deploy/push_desktop.sh --no-build`, or install it
manually). From 0.1.2 on, updates come from Gitea.
`package-type` in the app resources tells electron-updater whether to run
`pacman -U` or `dpkg -i` (both via pkexec/sudo); the per-user NSIS install
updates without elevation.
+5 -4
View File
@@ -2,12 +2,13 @@ appId: io.j621.desktop
productName: J621
copyright: Copyright (c) 2026 JakeBreath — Jake Labs Non-Commercial Software Licence
# Update feed served by the frontend nginx (deploy/data/desktop, published
# with deploy/push_desktop.sh). Baked into resources/app-update.yml; override
# at runtime with J621_UPDATE_URL for a fork or a test feed.
# Update feed: desktop/src/main.ts resolves the newest desktop-v* release on
# Gitea at check time (J621_UPDATE_REPO). This block only tells electron-builder
# to emit latest.yml/latest-linux.yml next to the installers; J621_UPDATE_URL
# overrides the feed for a fork or a test.
publish:
provider: generic
url: https://j621.rainbow-herring.ts.net/desktop
url: https://gitea.rainbow-herring.ts.net/JakeBreath/J621/releases
directories:
output: release
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "j621-desktop",
"version": "0.1.1",
"version": "0.1.2",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "j621-desktop",
"version": "0.1.1",
"version": "0.1.2",
"license": "LicenseRef-Jake-Labs-Non-Commercial",
"dependencies": {
"electron-updater": "^6.8.9"
+1 -1
View File
@@ -1,7 +1,7 @@
{
"name": "j621-desktop",
"productName": "J621",
"version": "0.1.1",
"version": "0.1.2",
"private": true,
"description": "Desktop shell for the J621 self-hosted media archive",
"author": {
+82 -4
View File
@@ -238,8 +238,9 @@ function isDownloadNavigation(url: URL): boolean {
/**
* Updates are manual by design: Linux packages install through pacman/dpkg
* (pkexec/sudo) and the Windows build is unsigned, so the app asks before
* downloading and again before installing. The feed comes from `publish` in
* electron-builder.yml and can be overridden with J621_UPDATE_URL.
* downloading and again before installing. The feed is resolved at check time
* from the newest `desktop-v*` release on Gitea (`J621_UPDATE_REPO`);
* `J621_UPDATE_URL` overrides it for forks and the smoke test.
*/
type UpdateEvent =
| { state: "available"; version: string }
@@ -250,9 +251,84 @@ type UpdateEvent =
let updateReporter: ((event: UpdateEvent) => void) | null = null;
let updateCheckRunning = false;
function setUpUpdates(win: BrowserWindow): void {
const UPDATE_REPO =
process.env.J621_UPDATE_REPO?.trim() ||
"https://gitea.rainbow-herring.ts.net/jakebreath/j621";
interface ReleaseAsset {
name: string;
browser_download_url: string;
}
interface Release {
tag_name: string;
draft: boolean;
prerelease: boolean;
assets?: ReleaseAsset[];
}
function releaseVersion(tag: string): [number, number, number] | null {
const match = /^desktop-v(\d+)\.(\d+)\.(\d+)$/.exec(tag);
if (!match) return null;
return [Number(match[1]), Number(match[2]), Number(match[3])];
}
function compareVersions(
a: [number, number, number],
b: [number, number, number],
): number {
for (let index = 0; index < 3; index += 1) {
if (a[index] !== b[index]) return a[index] - b[index];
}
return 0;
}
/**
* Resolve the generic feed base electron-updater should use.
*
* Gitea's API path is case-sensitive (owner/repo must match the login), while
* the asset URLs it returns are canonical, so the base is derived from the
* platform's metadata asset (`latest-linux.yml` / `latest.yml`).
*/
async function resolveReleaseFeed(): Promise<string> {
const override = process.env.J621_UPDATE_URL?.trim();
if (override) autoUpdater.setFeedURL({ provider: "generic", url: override });
if (override) return override;
const match = /^(https?:\/\/[^/]+)\/([^/]+)\/([^/]+?)\/?$/.exec(UPDATE_REPO);
if (!match) {
throw new Error(
`J621_UPDATE_REPO must be <origin>/<owner>/<repo> (got ${UPDATE_REPO}).`,
);
}
const [, origin, owner, repo] = match;
const response = await fetch(
`${origin}/api/v1/repos/${owner}/${repo}/releases?limit=50`,
{ headers: { Accept: "application/json" } },
);
if (!response.ok) {
throw new Error(`Release lookup on Gitea failed (HTTP ${response.status}).`);
}
const releases = (await response.json()) as Release[];
const assetName =
process.platform === "win32" ? "latest.yml" : "latest-linux.yml";
let best: { version: [number, number, number]; asset: ReleaseAsset } | null =
null;
for (const release of releases) {
if (release.draft || release.prerelease) continue;
const version = releaseVersion(release.tag_name);
if (!version) continue;
const asset = release.assets?.find((entry) => entry.name === assetName);
if (!asset) continue;
if (!best || compareVersions(version, best.version) > 0) {
best = { version, asset };
}
}
if (!best) {
throw new Error(`No ${assetName} asset found in ${UPDATE_REPO} releases.`);
}
return best.asset.browser_download_url.replace(/\/[^/]*$/, "");
}
function setUpUpdates(win: BrowserWindow): void {
if (!app.isPackaged) autoUpdater.forceDevUpdateConfig = true;
autoUpdater.autoDownload = false;
autoUpdater.autoInstallOnAppQuit = false;
@@ -336,6 +412,8 @@ async function checkForUpdates(win: BrowserWindow): Promise<void> {
if (updateCheckRunning) return;
updateCheckRunning = true;
try {
const feed = await resolveReleaseFeed();
autoUpdater.setFeedURL({ provider: "generic", url: feed });
await autoUpdater.checkForUpdates();
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
+3 -1
View File
@@ -15,6 +15,7 @@ import { ConfirmDialog } from "@/components/ConfirmDialog";
import { StatusFooter } from "@/components/StatusFooter";
import { StatusPill } from "@/components/StatusPill";
import { Toasts } from "@/components/Toasts";
import { UploadIndicator } from "@/components/UploadIndicator";
import { api } from "@/lib/api";
import { hasBackend } from "@/lib/backend";
import { cn } from "@/lib/cn";
@@ -112,7 +113,7 @@ export function AppShell() {
return (
<div className="flex min-h-screen flex-col">
<header className="sticky top-0 z-40 border-b border-ctp-surface0 bg-ctp-crust/95 backdrop-blur">
<div className="mx-auto flex h-14 w-full max-w-[1600px] items-center gap-4 px-4">
<div className="flex h-14 w-full items-center gap-4 px-3 sm:px-4">
<Link to="/" className="flex shrink-0 items-center">
<span className="rounded bg-ctp-mauve px-2 py-1 font-mono text-xs font-bold tracking-wide text-ctp-crust">
J621
@@ -153,6 +154,7 @@ export function AppShell() {
<div className="ml-auto flex items-center gap-3">
{backend ? (
<>
<UploadIndicator />
<StatusPill status={status} />
{user ? (
<>
+3 -3
View File
@@ -18,9 +18,9 @@ const ratingLabels: Record<string, string> = {
};
export function MediaCard({ item }: { item: MediaItem }) {
const preview = apiUrl(
item.kind === "video" ? item.thumbnail_url : item.raw_url,
);
// The thumbnail endpoint now builds real 480px previews for images too, so
// the grid no longer pulls full-size originals.
const preview = apiUrl(item.thumbnail_url);
const rating = item.display_rating;
return (
@@ -0,0 +1,56 @@
import { Link } from "react-router-dom";
import { cn } from "@/lib/cn";
import { useUploadStatus } from "@/lib/uploadStatus";
const PHASE_LABELS: Record<string, string> = {
md5: "e621 MD5",
visual: "Visual similarity",
iqdb: "IQDB",
};
/**
* Shell-wide progress for the server-side upload pipeline.
*
* Staged uploads keep processing after the Upload page is closed, so this
* pill is the "leave and keep an eye on it" view: it lives in the header on
* every page and links back to the board.
*/
export function UploadIndicator() {
const { data: status } = useUploadStatus();
if (!status || !status.active) return null;
const total = Math.max(status.total, status.processed + status.failed);
const done = status.processed + status.failed;
const percent = total > 0 ? Math.min(100, Math.round((done / total) * 100)) : 0;
const broken = status.status === "paused" || status.status === "error";
const label =
status.status === "paused"
? "uploads paused"
: status.status === "error"
? "uploads failed"
: (PHASE_LABELS[status.phase] ?? "processing uploads");
const count = total > 0 ? `${done}/${total}` : `${status.outstanding} left`;
return (
<Link
to="/upload"
title={status.error || "Staged uploads are being processed in the background"}
className={cn(
"flex items-center gap-2 rounded-md border px-2 py-1 font-mono text-[10px] transition",
broken
? "border-ctp-red/40 text-ctp-red hover:bg-ctp-red/10"
: "border-ctp-surface1 text-ctp-subtext0 hover:bg-ctp-surface0 hover:text-ctp-text",
)}
>
<span className="hidden sm:inline">{label}</span>
<span>{count}</span>
<span className="h-1 w-12 overflow-hidden rounded-full bg-ctp-surface0">
<span
className={cn("block h-full", broken ? "bg-ctp-red" : "bg-ctp-teal")}
style={{ width: `${percent}%` }}
/>
</span>
</Link>
);
}
+29 -12
View File
@@ -1,5 +1,5 @@
import { useMutation, useQuery, useQueryClient } from "@tanstack/react-query";
import { Eye, StarOff } from "lucide-react";
import { ChevronDown, ChevronRight, Eye, StarOff } from "lucide-react";
import { useState } from "react";
import { Link } from "react-router-dom";
@@ -148,28 +148,43 @@ function BlacklistCloudPanel() {
q.state.data?.status === "building" ? 2000 : false,
});
const cloud = query.data;
const [expanded, setExpanded] = useState(false);
return (
<section className="rounded-lg border border-ctp-surface0 bg-ctp-base p-4">
<div className="flex flex-wrap items-center justify-between gap-2">
<button
type="button"
onClick={() => setExpanded((value) => !value)}
title={expanded ? "Hide the tags" : "Show the tags"}
className="flex w-full items-center justify-between gap-2 text-left"
>
<h2 className="text-sm font-semibold text-ctp-subtext1">
Blacklisted tags
</h2>
<p className="text-xs text-ctp-overlay0">
<span className="flex items-center gap-1.5 font-mono text-[11px] text-ctp-overlay0">
{cloud ? `${cloud.blacklist_count} entries` : "…"}
{cloud?.status === "building" ? (
<Spinner className="h-3 w-3" />
) : null}
{expanded ? (
<ChevronDown className="h-3.5 w-3.5" />
) : (
<ChevronRight className="h-3.5 w-3.5" />
)}
</span>
</button>
{expanded ? (
<>
<p className="mt-2 text-xs text-ctp-overlay0">
{cloud?.source === "user"
? "From your e621 blacklist"
: "From e621's anonymous default blacklist"}
{cloud ? ` · ${cloud.blacklist_count} entries` : ""}
{cloud ? ` · ${cloud.posts} feed post(s) scanned` : ""}
{cloud?.computed_at ? ` · computed ${formatDate(cloud.computed_at)}` : ""}
{cloud?.status === "building" ? (
<span className="ml-2 inline-flex items-center gap-1.5">
<Spinner className="h-3 w-3" /> building…
</span>
) : null}
{cloud?.computed_at
? ` · computed ${formatDate(cloud.computed_at)}`
: ""}
</p>
</div>
{query.isPending ? (
<div className="mt-3 flex justify-center py-4">
<Spinner className="h-4 w-4" />
@@ -190,6 +205,8 @@ function BlacklistCloudPanel() {
No blacklisted tags in your feeds.
</p>
)}
</>
) : null}
</section>
);
}
+22 -3
View File
@@ -1,5 +1,5 @@
import { keepPreviousData, useMutation, useQuery, useQueryClient } from "@tanstack/react-query";
import { X } from "lucide-react";
import { ChevronDown, ChevronRight, X } from "lucide-react";
import { useEffect, useMemo, useRef, useState } from "react";
import { Link, useLocation, useSearchParams } from "react-router-dom";
@@ -43,6 +43,7 @@ export default function OnlinePage() {
const queryClient = useQueryClient();
const [draft, setDraft] = useState<string | null>(null);
const [blacklistDraft, setBlacklistDraft] = useState("");
const [showBlacklist, setShowBlacklist] = useState(false);
const tagInput = draft ?? tags;
const zoom = useUi((state) => state.zoom);
const perPage = useUi((state) => state.e621PerPage);
@@ -271,10 +272,27 @@ export default function OnlinePage() {
</div>
<div className="flex flex-col gap-2">
<button
type="button"
onClick={() => setShowBlacklist((value) => !value)}
title={showBlacklist ? "Hide the blacklist" : "Show the blacklist"}
className="flex items-center justify-between gap-2 text-left"
>
<span className="text-xs font-medium uppercase tracking-wide text-ctp-overlay1">
Your blacklist
</span>
{credentials?.configured ? (
<span className="flex items-center gap-1 font-mono text-[11px] text-ctp-overlay0">
{blacklistEntries.length} tag
{blacklistEntries.length === 1 ? "" : "s"}
{showBlacklist ? (
<ChevronDown className="h-3.5 w-3.5" />
) : (
<ChevronRight className="h-3.5 w-3.5" />
)}
</span>
</button>
{showBlacklist ? (
credentials?.configured ? (
<>
{blacklistEntries.length > 0 ? (
<div className="flex flex-wrap gap-1.5">
@@ -337,7 +355,8 @@ export default function OnlinePage() {
</Link>{" "}
to manage your blacklist.
</p>
)}
)
) : null}
</div>
<div className="flex flex-col gap-2">
File diff suppressed because it is too large Load Diff
+12 -9
View File
@@ -89,18 +89,19 @@ export function effectiveCredentials(
const GIT_HASH = typeof __GIT_HASH__ === "string" ? __GIT_HASH__ : "dev";
const CLIENT_VERSION = `J621/${GIT_HASH} (JakeBreath)`;
// e621 allows 2 requests/second hard, 1/second sustained — and the IQDB
// endpoint is stricter, so stay comfortably under it. Serialize every request
// through a queue with a minimum gap.
// e621 allows 2 requests/second hard, 1/second sustained. Serialize every
// request through a queue with a minimum gap; IQDB is throttled much harder
// by e621, so it gets a wider gap of its own.
let lastRequestAt = 0;
let queue: Promise<unknown> = Promise.resolve();
/** A hung request would block the whole serialized queue forever. */
const REQUEST_TIMEOUT_MS = 20_000;
const REQUEST_GAP_MS = 1500;
const REQUEST_GAP_MS = 1000;
const IQDB_GAP_MS = 2500;
/** A 429 (or a CORS-blocked failure) pauses every e621 call for a while. */
const RATE_LIMIT_COOLDOWN_MS = 60_000;
/** A 429 (or a CORS-blocked failure) pauses every e621 call briefly. */
const RATE_LIMIT_COOLDOWN_MS = 15_000;
const COOLDOWN_KEY = "j621.e621.cooldown";
let cooldownUntil = 0;
@@ -143,10 +144,10 @@ function schedule<T>(task: () => Promise<T>): Promise<T> {
return run;
}
async function throttle(): Promise<void> {
async function throttle(minGap = REQUEST_GAP_MS): Promise<void> {
const wait = Math.max(
0,
lastRequestAt + REQUEST_GAP_MS - Date.now(),
lastRequestAt + minGap - Date.now(),
e621CooldownRemainingMs(),
);
if (wait > 0) {
@@ -168,7 +169,9 @@ export function e621Request<T>(
options: E621RequestOptions = {},
): Promise<T> {
return schedule(async () => {
await throttle();
await throttle(
path.includes("iqdb_queries") ? IQDB_GAP_MS : REQUEST_GAP_MS,
);
const base = credentials.base_url.replace(/\/+$/, "");
const url = new URL(`${base}/${path.replace(/^\/+/, "")}`);
+57 -5
View File
@@ -355,12 +355,17 @@ export interface TempUpload {
status: "pending" | "visual_match" | "completed" | "error";
resolution: "" | "auto_md5" | "duplicate" | "linked" | "custom";
e621_post_id: number | null;
e621_data: E621StoredPost | null;
/** Full post payload: only present on the detail endpoint. */
e621_data?: E621StoredPost | null;
custom_rating: Rating;
custom_tags: string[];
custom_notes: string;
iqdb_data: E621IqdbCandidate[] | null;
visual_matches:
/** Only present on the detail endpoint. */
custom_tags?: string[];
/** Only present on the detail endpoint. */
custom_notes?: string;
/** Only present on the detail endpoint. */
iqdb_data?: E621IqdbCandidate[] | null;
/** Only present on the detail endpoint. */
visual_matches?:
| {
j_id: string;
filename: string;
@@ -371,10 +376,57 @@ export interface TempUpload {
library_j_id: string | null;
file_url: string | null;
preview_url: string | null;
/** Background pipeline state (see the UploadRun status endpoint). */
pipeline_error: string;
attempts: number;
md5_checked?: boolean;
visual_checked?: boolean;
iqdb_checked?: boolean;
processing?: boolean;
similar_count?: number;
created_at: string;
updated_at: string;
}
/** One upload the pipeline indexed while the page was watching. */
export interface UploadCompletion {
id: string;
filename: string;
j_id: string;
resolution: string;
post_id: number | null;
thumbnail_url: string | null;
at: string;
}
/** Response of POST /api/uploads/ — a pending row or an instant completion. */
export interface UploadStagingResult {
temp_id: string;
original_filename: string;
status: "pending" | "visual_match" | "completed" | "error";
resolution: string;
library_j_id: string | null;
preview_url: string | null;
e621_post_id?: number | null;
}
/** Progress of the server-side staged-upload pipeline. */
export interface UploadStatus {
status: "idle" | "running" | "paused" | "error";
active: boolean;
phase: "" | "md5" | "visual" | "iqdb";
total: number;
processed: number;
matched: number;
failed: number;
error: string;
outstanding: number;
waiting: { md5: number; visual: number; iqdb: number };
/** Session notification feed: indexed uploads, newest first. */
recent_completions: UploadCompletion[];
updated_at: string | null;
}
export interface StatJob {
kind: "download" | "match";
id: string;
+30
View File
@@ -0,0 +1,30 @@
import { useQuery } from "@tanstack/react-query";
import { api } from "@/lib/api";
import { hasBackend } from "@/lib/backend";
import type { UploadStatus } from "@/lib/types";
import { useAuth } from "@/store/auth";
export const uploadStatusQueryKey = ["upload-status"] as const;
/**
* Poll the server-side upload pipeline while it has work and stop once it is
* idle. Shared by the Upload page and the shell indicator so both show the
* same state without extra requests.
*/
export function useUploadStatus(enabled = true) {
const user = useAuth((state) => state.user);
return useQuery({
queryKey: uploadStatusQueryKey,
queryFn: () => api<UploadStatus>("/api/uploads/status/"),
enabled: enabled && hasBackend() && Boolean(user?.can_upload),
// TanStack takes the smallest interval across the observers, so the
// Upload page and the shell pill can safely share this query.
refetchInterval: (query) => (query.state.data?.active ? 2_000 : false),
});
}
/** Ask the server to (re)start processing the caller's staged uploads. */
export function postProcessUploads(): Promise<UploadStatus> {
return api<UploadStatus>("/api/uploads/process/", { method: "POST" });
}