Files
rnes/STATE.md
T
2026-08-09 16:53:53 -05:00

12 KiB

Status

Done

  • Repo set up, backed up the original game ROMs to roms/backup (gitignored, stays local)
  • Got a full set of CPU test ROMs in roms/test/cpu: kevtris nestest (with reference log) plus the blargg suite (instr tests, timing, interrupts, dummy reads, reset checks)
  • Cartridge module done and working (src/cartridge.rs):
    • iNES header parser (magic, PRG/CHR sizes, mapper, mirroring, flags)
    • Mapper 0 (NROM) support only: 16KB PRG mirroring, 32KB direct, CHR-ROM and CHR-RAM (zeroed 8KB buffer)
    • read_prg/read_chr/write_chr with debug_assert guards
    • Split from_file (thin I/O wrapper) / parse_data (pure parser) for testability
    • 8 unit tests, all passing (header, PRG mapping, error cases, CHR-RAM roundtrip)
  • Bus module done (src/bus.rs):
    • Bus trait (read/write) so the CPU is decoupled from the concrete memory layout
    • NesBus: 2KB mirrored CPU RAM, cartridge PRG routing, PPU/APU/SRAM/expansion regions stubbed (return 0)
  • CPU skeleton done (src/cpu.rs):
    • Flag bit constants, registers (A/X/Y/SP/PC/P), set/clear/flag helpers, set_zp_flags
    • Memory helpers: fetch, fetch16, read, write, read16, push, pop
    • All 10 addressing mode helpers, including the JMP ($xxxx) page-wrap bug and page-cross cycle penalties
    • First 3 opcodes implemented (JMP abs, LDX imm, STX zp)
    • step dispatch shell + trace in nestest format
    • set_test_start for the nestest harness
  • nestest harness working (main.rs): loads nestest.nes, runs from $C000 with CYC:7, dumps trace to ours.log
  • Batch 1 done (subroutines + stack): JSR, RTS, PHA, PLA, PHP, PLP, RTI, BRK, plus NOP and SEC as quick wins. Verified SP:FD->FB across JSR (push high-then-low works)
  • Batch 2 done (branches): all 8 branches (BEQ/BNE/BCC/BCS/BPL/BMI/BVC/BVS) via one shared branch() routine with condition param. Taken = +3 cycles, page-cross = +1 more. Verified both taken and not-taken paths in the log
  • Flag set/clear ops done (part of Batch 10): CLC, CLI, CLV, CLD, SEI, SED (SEC done in Batch 1). SEI/SED set, the rest clear. Fixed a real bug where SEI/SED were clearing instead of setting
  • Batch 8 done (inc/dec): INC (4 modes), DEC (4 modes), INX, INY, DEX, DEY. RMW cores (inc/dec) read-modify-write and correctly return base WITHOUT the page-cross extra (RMW cycles are fixed). Register ops are wrapping_add/sub with Z/N only. 124 of 151 official opcodes done. INY/INX carried the log from 506 to 678 lines
  • Batch 9 done (shifts): ASL, LSR, ROL, ROR (accumulator + 4 memory modes each) via four shift cores (shift_asl/lsr/rol/ror) + wrappers. ROL/ROR read carry-in before overwriting it. Memory shifts follow the RMW fixed-cycle rule. 144 of 151
  • Batch 10 done (transfers + JMP ind): TAX, TAY, TXA, TYA, TSX, TXS (all 2 cycles; TXS sets NO flags, the other five set Z/N), plus JMP (0x6C) which finally uses address_ind. Resolved the long-standing address_ind dead-code warning
  • OFFICIAL CPU COMPLETE: all 151 official opcodes implemented. Trace matches the nestest reference log through line 5003 (entire official section), 0 mismatches. Stops at line 5004 (0x04, first illegal opcode) by design
  • Fixed a latent Batch 4 bug the log caught at line 3319: the sta core added the page-cross extra, but stores have FIXED cycle counts. Fixed sta core to return base without extra, and corrected bases (sta_abx/sta_aby 5, sta_izy 6). Before the fix the diff had a +1 CYC drift from line 3319 onward
  • Batch 3 done (loads): LDA (8 modes), LDX (5 modes), LDY (5 modes) via three shared cores (lda/ldx/ldy) + thin wrappers. Caught two real bugs in review: all 4 LDY wrappers called self.ldx (would write to X instead of Y), and 6 of 8 LDX/LDY wrappers had wrong base cycles (3 instead of 4). Both fixed
  • Batch 4 done (stores): STA (7 modes), STX (3 modes), STY (3 modes) via three shared cores (sta/stx/sty) + thin wrappers. No flag changes. Refactored stx_zp to use the shared stx core for consistency. Fixed a comment typo (sta_abx was labeled 0x95 instead of 0x9D)
  • Batch 5 done (compares): CMP (8 modes), CPX (3 modes), CPY (3 modes) via three shared cores (cmp/cpx/cpy) + wrappers + immediate variants. Sets C (reg >= mem), Z, N; no store. Uses wrapping_sub for the borrow math. 70 of 151 official opcodes done. No diff progress expected until the log reaches these opcodes
  • Batch 6 done (arithmetic): ADC (8 modes), SBC (8 modes), BIT (2 modes) via shared add_with_carry core + adc/sbc/bit cores + wrappers. SBC adds the one's complement (A + ~mem + C), which makes the V-flag math correct. Decimal flag inert (binary only). 88 of 151 official opcodes done. Big diff progress: BIT/PHP/PLA/PHA carried the log from 38 to 73 lines
  • Batch 7 done (logic): AND (8 modes), ORA (8 modes), EOR (8 modes) via three cores (and/ora/eor) + wrappers + immediate variants. Only Z/N flags affected, C/V untouched. Caught a real bug in review: all three _izx wrappers used base 5 instead of 6. 112 of 151 official opcodes done. Huge diff progress: carried the log from 73 to 506 lines
  • Match compaction done: step and opcode_len both reformatted to max 4 opcodes per line. step match grouped by family with header comments; opcode_len keeps per-entry /* */ mnemonics. Diff verified unchanged after compaction
  • trace bytes are length-aware via opcode_len (fixed-width byte field, columns align for awk diff)
  • MMC1 Phase 1 done (src/mapper.rs): Mirroring enum (5 variants incl. OneScreenLow/High), trait Mapper: Debug, Nrom impl, Mmc1 impl (shift-register protocol, PRG 32K/16K banking, CHR 8K/two-4K banking, init control=0x0C)
  • MMC1 Phase 2 done (src/cartridge.rs rewired): mapper is Box, 8KB prg_ram added, derives now Debug only, from_file accepts mapper 0/1, read/write delegate to mapper, mirroring() accessor + read/write_prg_ram
  • MMC1 Phase 3 done (src/bus.rs): $6000-$7FFF routes to cart prg_ram, $8000-$FFFF writes route to cart.write_prg
  • MMC1 Phase 4 done (regression gate): nestest still matches 5004 lines / 0 mismatches, builds and runs clean
  • MMC1 Phase 5 partial (blargg harness in main.rs): cargo run = nestest mode; cargo run -- = blargg mode (cpu.reset, cycle cap, reports $6000-$6003). official_only.nes returns $6000=80 = blargg "test in progress" sentinel - the test waits for NMI (vblank) which we cannot produce without a PPU. Deferred until PPU/APU exist (blargg suite is the LAST milestone anyway)
  • NMI/IRQ servicing done (src/cpu.rs): set_nmi/set_irq setters, shared interrupt() helper (push PC hi/lo, push p with B CLEAR unlike BRK, set I, jump vector, 7 cycles), NMI checked before IRQ at the top of step, IRQ masked by the I flag, cycles accumulated on both interrupt paths. cpu_interrupts.nes currently reports $6000=01 - it runs but needs the APU frame-counter IRQ source ($4017) to pass, which is roadmap Phase 6; verification deferred to the last milestone
  • Backup library mapper audit: mappers 0 (NROM) and 1 (MMC1) cover many games; still need mapper 2 (UNROM: Castlevania, Contra, Megaman 1), mapper 4 (MMC3: SMB2, SMB3, Lolo 2), mapper 7 (AOROM: Who Framed Roger Rabbit)
  • Finding: real commercial games rarely use illegal opcodes; none of the backup library needs them. Official-only CPU is sufficient for the goal of playing these games

Next

REWEIGHTED ROADMAP: priority is now building the rest of the NES hardware so games actually run visibly, instead of chasing the full blargg/illegal-opcode pass. Goal: see the backup library running in a GUI. First test game: SMB (NROM, no new mapper needed). Blargg full suite + illegal opcodes are LAST, once the hardware can display test output on the NES screen.

Phases in order:

  1. NMI/IRQ servicing - DONE (see Done section). cpu_interrupts.nes verification deferred to the last milestone (needs APU frame-counter IRQ)

  2. PPU-1 rendering core (new src/ppu.rs): registers $2000-$2007 routed in bus (currently stubbed 0), VRAM (nametables, pattern tables, palettes), background rendering into 256x240 framebuffer of palette indices, $2002 status with vblank flag, mirroring from cart, vblank raises NMI. Acceptance: SMB title screen renders, verified headless by framebuffer dump. NEXT. Confirmed decisions: full dot-accurate scanline model (sprites in PPU-2 need cycle timing, MMC3 needs scanline IRQ), scroll registers $2005/$2006 correct now (v/t/x/w loopy system), Cartridge passed as a parameter to read/write/tick (no Rc/RefCell). Build steps:

    2a. DONE - palette.rs (64-color 2C02 SYSTEM_PALETTE in index order + index_to_rgb with & 0x3F mask) and render.rs (trait Renderer with present(framebuffer, width, height) + PpmRenderer + pure ppm_bytes for testability). 4 unit tests pass (known palette values, index masking, PPM header, per-pixel RGB routing). Note: 2C02 chosen over composite palettes - we render the raw PPU output. src/ppu.rs exists as an empty placeholder 2b. Ppu struct + registers $2000-$2007 + VRAM routing ($0000-$1FFF = cart CHR via cart.read_chr/write_chr, $2000-$2FFF = nametables in vram[0x800] with cart.mirroring(), $3F00-$3FFF = palette[32] with $3F10-$3F1F mirror) + bus wiring ($2000-$3FFF reads/writes via disjoint-field borrows). Verify with a debug harness writing $2006/$2007 and reading back. NEXT 2c. Frame timing + vblank + NMI: 341 dots/scanline, 262 scanlines/frame, 3 PPU dots per CPU cycle. tick(cycles) advances cycles*3 dots. Scanline 241 dot 0 sets vblank flag + nmi_pending. Frame loop bridges: take_nmi() -> cpu.set_nmi(). Verify $2002 flag toggles and NMI fires 2d. Static background render at scroll 0: nametable -> attribute -> pattern -> palette into 256x240 framebuffer. May show messy SMB output before scroll lands - expected 2e. Scroll registers v/t/x/w (loopy system) + the dot-by-dot background pipeline: shift registers (bg_pattern_lo/hi 16-bit, bg_attr 16-bit, tile_latch, attr_latch), 4 fetch stages per 8-dot group (nametable, attribute, pattern low, pattern high), coarse X increment every 8 dots with nametable flip, coarse Y/fine Y wrap at scanline end. The densest code in the project 2f. Acceptance: SMB title screen renders recognizable, dumped to PPM

  3. PPU-2 sprites + scrolling + $4014 OAM DMA: makes SMB actually playable

  4. GUI (egui + winit, first external deps): display framebuffer, 60fps loop, keyboard -> controller ($4016/$4017). First "games running on screen" moment

  5. Mappers: 2 (UNROM), 7 (AOROM) simple; 4 (MMC3) complex - includes scanline IRQ counter required for SMB3 status bar. Verified by running the real backups

  6. APU audio: 5 channels, frame counter -> IRQ, mixing. Verified with SMB

  7. LAST: blargg full suite (screen output becomes observable) + illegal-opcode phase for the full 8991-line nestest pass

Design note for PPU: PPU takes &mut Cartridge as a parameter to read/write/tick (CHR via cart.read_chr/write_chr, mirroring via cart.mirroring). Frame loop owns bus + cpu + renderer and drives all three - loop bridges NMI via ppu.take_nmi() -> cpu.set_nmi(), so neither PPU nor bus holds a &mut Cpu. Framebuffer = palette indices (1 byte/pixel), RGB conversion happens in the Renderer (PPM now, GUI later).

Blargg/illegal background: ~76 illegal opcodes in the nestest log (23 undocumented NOPs, SLO/RLA/SRE/RRA, LAX/SAX, DCP/ISC, EB=SBC alias). Reference: nesdev wiki. Not needed for the backup library but wanted for the full log pass.

Decisions

  • No external dependencies yet, keeping it pure std until the GUI milestone
  • GUI choice: egui + winit (decided, not yet used)
  • Scope: cartridge + CPU first, verified headless before graphics
  • Bus is a trait (impl Bus) so the CPU works against any memory layout (real bus, test bus, debug bus)
  • Flags stored as a raw u8 with named bit-mask constants (FLAG_CARRY etc.)
  • opcode length is a growing match in opcode_len() (returns u8); migrate to a full 256-entry table when it gets big
  • trace diff compares PC + A/X/Y/P/SP/CYC, ignoring disassembly/PPU columns
  • Mapper 0 first, other mappers (MMC1, CNROM) implemented after CPU works
  • Mappers use a trait-based design (Mapper trait + Box) so new mappers are self-contained impl blocks; ROM data owned by Cartridge and passed in per call
  • PRG RAM lives on the Cartridge (8KB), the bus interconnects it
  • MMC1 (mapper 1) done for the cartridge; mappers 2/4/7 planned after PPU
  • NMI/IRQ servicing done; APU frame counter (IRQ source for cpu_interrupts.nes) comes with the APU milestone
  • Priority reweighted: build the rest of the NES hardware (PPU -> GUI -> mappers -> APU) so games run visibly; blargg full suite + illegal opcodes are last
  • Framebuffer is palette indices (1 byte/pixel) from the PPU; RGB conversion happens in the GUI
  • Test ROMs are homebrew/open, no ethics issue; game backups stay in gitignored roms/backup