diff --git a/.gitignore b/.gitignore index 2133345..2b19309 100644 --- a/.gitignore +++ b/.gitignore @@ -3,3 +3,4 @@ roms/ AGENTS/ AGENTS.md Cargo.lock +renders/ diff --git a/STATE.md b/STATE.md index 2b81d6d..a6eb319 100644 --- a/STATE.md +++ b/STATE.md @@ -42,6 +42,8 @@ - MMC1 Phase 4 done (regression gate): nestest still matches 5004 lines / 0 mismatches, builds and runs clean - MMC1 Phase 5 partial (blargg harness in main.rs): cargo run = nestest mode; cargo run -- = blargg mode (cpu.reset, cycle cap, reports $6000-$6003). official_only.nes returns $6000=80 = blargg "test in progress" sentinel - the test waits for NMI (vblank) which we cannot produce without a PPU. Deferred until PPU/APU exist (blargg suite is the LAST milestone anyway) - NMI/IRQ servicing done (src/cpu.rs): set_nmi/set_irq setters, shared interrupt() helper (push PC hi/lo, push p with B CLEAR unlike BRK, set I, jump vector, 7 cycles), NMI checked before IRQ at the top of step, IRQ masked by the I flag, cycles accumulated on both interrupt paths. cpu_interrupts.nes currently reports $6000=01 - it runs but needs the APU frame-counter IRQ source ($4017) to pass, which is roadmap Phase 6; verification deferred to the last milestone +- PPU 2a done (src/palette.rs + src/render.rs): 64-color 2C02 SYSTEM_PALETTE in index order, index_to_rgb with & 0x3F mask. Renderer trait (the only hardware trait) + PpmRenderer writing numbered frames to renders/ (frame_0000.ppm etc.) + pure ppm_bytes for testability. 4 unit tests pass. 2C02 chosen over composite palettes - we render raw PPU output +- PPU 2b done (src/ppu.rs struct + registers + VRAM routing + src/bus.rs open-bus): Ppu has ctrl/mask/status/oam/oamaddr, shared write_latch (W bit for $2005/$2006), vram_addr (14-bit), data_buffer ($2007 read-buffer), vram[0x1000] (4KB for FourScreen), palette[32]. read/write handle $2000-$2007 with open-bus for write-only regs ($2000/$2001/$2003/$2005/$2006), $2002 read = (status & 0xE0) | (open_bus & 0x1F) + clears vblank + resets write latch, $2004 write increments oamaddr, $2007 palette reads bypass the read-buffer. VRAM routing: $0000-$1FFF = CHR via cart, $2000-$3EFF = nametables via cart.mirroring() (H/V/OneScreenLow/High/FourScreen), $3F00-$3FFF = palette with 10/14/18/1C mirror. Bus tracks open_bus (last byte on CPU data bus, updated on every read/write) and passes it into ppu.read; PPU regs routed via disjoint-field borrows; APU/expansion stubs now return open_bus instead of 0. cartridge.rs mirroring honors the four-screen header bit - Backup library mapper audit: mappers 0 (NROM) and 1 (MMC1) cover many games; still need mapper 2 (UNROM: Castlevania, Contra, Megaman 1), mapper 4 (MMC3: SMB2, SMB3, Lolo 2), mapper 7 (AOROM: Who Framed Roger Rabbit) - Finding: real commercial games rarely use illegal opcodes; none of the backup library needs them. Official-only CPU is sufficient for the goal of playing these games @@ -52,21 +54,22 @@ REWEIGHTED ROADMAP: priority is now building the rest of the NES hardware so gam Phases in order: 1. NMI/IRQ servicing - DONE (see Done section). cpu_interrupts.nes verification deferred to the last milestone (needs APU frame-counter IRQ) -2. PPU-1 rendering core (new src/ppu.rs): registers $2000-$2007 routed in bus (currently stubbed 0), VRAM (nametables, pattern tables, palettes), background rendering into 256x240 framebuffer of palette indices, $2002 status with vblank flag, mirroring from cart, vblank raises NMI. Acceptance: SMB title screen renders, verified headless by framebuffer dump. NEXT. Confirmed decisions: full dot-accurate scanline model (sprites in PPU-2 need cycle timing, MMC3 needs scanline IRQ), scroll registers $2005/$2006 correct now (v/t/x/w loopy system), Cartridge passed as a parameter to read/write/tick (no Rc/RefCell). Build steps: +2. PPU-1 rendering core - PARTIAL: 2a (palette/render) and 2b (registers/VRAM/open-bus) DONE (see Done section). Remaining: full background rendering in ONE merged step - timing + scroll + background pipeline together, no intermediate gating (we do not test until the full PPU exists). Acceptance: SMB title screen renders to PPM via main.rs run_cart. NEXT. Confirmed decisions: full dot-accurate scanline model, loopy v/t/x/w scroll system, Cartridge passed as parameter, open-bus tracking in the bus, hardware-accurate throughout (same philosophy as the CPU). Merged step: - 2a. DONE - palette.rs (64-color 2C02 SYSTEM_PALETTE in index order + index_to_rgb with & 0x3F mask) and render.rs (trait Renderer with present(framebuffer, width, height) + PpmRenderer + pure ppm_bytes for testability). 4 unit tests pass (known palette values, index masking, PPM header, per-pixel RGB routing). Note: 2C02 chosen over composite palettes - we render the raw PPU output. src/ppu.rs exists as an empty placeholder - 2b. Ppu struct + registers $2000-$2007 + VRAM routing ($0000-$1FFF = cart CHR via cart.read_chr/write_chr, $2000-$2FFF = nametables in vram[0x800] with cart.mirroring(), $3F00-$3FFF = palette[32] with $3F10-$3F1F mirror) + bus wiring ($2000-$3FFF reads/writes via disjoint-field borrows). Verify with a debug harness writing $2006/$2007 and reading back. NEXT - 2c. Frame timing + vblank + NMI: 341 dots/scanline, 262 scanlines/frame, 3 PPU dots per CPU cycle. tick(cycles) advances cycles*3 dots. Scanline 241 dot 0 sets vblank flag + nmi_pending. Frame loop bridges: take_nmi() -> cpu.set_nmi(). Verify $2002 flag toggles and NMI fires - 2d. Static background render at scroll 0: nametable -> attribute -> pattern -> palette into 256x240 framebuffer. May show messy SMB output before scroll lands - expected - 2e. Scroll registers v/t/x/w (loopy system) + the dot-by-dot background pipeline: shift registers (bg_pattern_lo/hi 16-bit, bg_attr 16-bit, tile_latch, attr_latch), 4 fetch stages per 8-dot group (nametable, attribute, pattern low, pattern high), coarse X increment every 8 dots with nametable flip, coarse Y/fine Y wrap at scanline end. The densest code in the project - 2f. Acceptance: SMB title screen renders recognizable, dumped to PPM + 2c. Full background renderer (merged timing + scroll + pipeline): + - Ppu gains: scanline (0-261), cycle (0-340), framebuffer[256*240] palette indices, nmi_pending, frame_complete, loopy registers t (temp addr) + x (fine X) + w (write latch), background pipeline state (bg_pattern_lo/hi u16 shift regs, bg_attr u16, tile_latch, attr_latch) + - Timing: tick(cart, cycles) advances cycle by cycles*3; cycle>=341 -> scanline++; odd-frame skip on pre-render line (261 is 340 dots on odd frames); scanline 241 dot 0 sets vblank + nmi_pending; scanline 261 dot 0 clears vblank; frame_complete at 261 end. Rendering on 0-239 and 261 (pre-render fetches, no output); 240 idle + - Loopy scroll: $2005 first: t=(t&0x7FE0)|(val>>3), x=val&7, w=true; $2005 second: t=(t&0x0C1F)|((val&0xF8)<<2)|((val&7)<<12), w=false; $2006 first: t=(t&0x00FF)|((val&0x3F)<<8), w=true; $2006 second: t=(t&0x7F00)|val, v=t, w=false; $2002 read resets w + - Background fetch pipeline (4 fetch / 8 dot cycle): dot%8==0 nametable byte -> tile_latch, ==2 attribute -> attr_latch, ==4 pattern lo from CHR, ==6 pattern hi, ==7 reload shift regs + increment coarse X (31->0 flips nametable X bit). Each dot outputs pixel combining bg_pattern_lo/hi top bit + attr bit for 2-bit palette select -> background palette -> framebuffer. Vertical increment at scanline end (fine Y 7->0, coarse Y 29->0 flips nametable Y bit) + - main.rs: add use crate::render::{PpmRenderer, Renderer}, define frames (e.g. 3) in run_cart, cargo run -- boots from reset and renders frames to renders/. run_cart already calls the real methods (tick/frame_done/take_nmi/framebuffer/begin_frame) + - Acceptance (single gate): cargo run -- roms/backup/SMB.nes -> renders/frame_0000.ppm recognizable SMB title screen 3. PPU-2 sprites + scrolling + $4014 OAM DMA: makes SMB actually playable 4. GUI (egui + winit, first external deps): display framebuffer, 60fps loop, keyboard -> controller ($4016/$4017). First "games running on screen" moment 5. Mappers: 2 (UNROM), 7 (AOROM) simple; 4 (MMC3) complex - includes scanline IRQ counter required for SMB3 status bar. Verified by running the real backups 6. APU audio: 5 channels, frame counter -> IRQ, mixing. Verified with SMB 7. LAST: blargg full suite (screen output becomes observable) + illegal-opcode phase for the full 8991-line nestest pass -Design note for PPU: PPU takes &mut Cartridge as a parameter to read/write/tick (CHR via cart.read_chr/write_chr, mirroring via cart.mirroring). Frame loop owns bus + cpu + renderer and drives all three - loop bridges NMI via ppu.take_nmi() -> cpu.set_nmi(), so neither PPU nor bus holds a &mut Cpu. Framebuffer = palette indices (1 byte/pixel), RGB conversion happens in the Renderer (PPM now, GUI later). +Design note for PPU: PPU takes &mut Cartridge as a parameter to read/write/tick (CHR via cart.read_chr/write_chr, mirroring via cart.mirroring). Frame loop owns bus + cpu + renderer and drives all three - loop bridges NMI via ppu.take_nmi() -> cpu.set_nmi(), so neither PPU nor bus holds a &mut Cpu. Framebuffer = palette indices (1 byte/pixel), RGB conversion happens in the Renderer (PPM now, GUI later). Open bus tracked in the bus (last byte on CPU data bus, updated on every read/write), passed into ppu.read; write-only PPU regs return it. Hardware accuracy throughout - no accuracy questions pending, we build it correct like the CPU. Blargg/illegal background: ~76 illegal opcodes in the nestest log (23 undocumented NOPs, SLO/RLA/SRE/RRA, LAX/SAX, DCP/ISC, EB=SBC alias). Reference: nesdev wiki. Not needed for the backup library but wanted for the full log pass. diff --git a/src/bus.rs b/src/bus.rs index 6eb2469..c1b4206 100644 --- a/src/bus.rs +++ b/src/bus.rs @@ -16,6 +16,22 @@ impl NesBus { pub fn new(cart: Cartridge) -> Self { Self { ram: [0; 0x800], cart, ppu: Ppu::new(), openbus: 0 } } + + pub fn tick_ppu(&mut self, cycles: u64) { + self.ppu.tick(&mut self.cart, cycles); + } + pub fn frame_done(&self) -> bool { + self.ppu.frame_done() + } + pub fn take_nmi(&mut self) -> bool { + self.ppu.take_nmi() + } + pub fn framebuffer(&self) -> &[u8] { + self.ppu.framebuffer() + } + pub fn begin_frame(&mut self) { + self.ppu.begin_frame(); + } } impl Bus for NesBus { diff --git a/src/main.rs b/src/main.rs index 8466f6c..c648ae5 100644 --- a/src/main.rs +++ b/src/main.rs @@ -9,6 +9,8 @@ mod ppu; use crate::bus::{Bus, NesBus}; use crate::cartridge::Cartridge; use crate::cpu::Cpu; +use crate::render::{PpmRenderer, Renderer}; + fn main() { let args: Vec = std::env::args().collect(); @@ -24,17 +26,18 @@ fn run_cart(path: &str) { let cart = Cartridge::from_file(path).unwrap(); let mut bus = NesBus::new(cart); let mut cpu = Cpu::new(); - cpu.reset(&mut bus); // boot from $FFFC-$FFFD + cpu.reset(&mut bus); let mut renderer = PpmRenderer::new(); + let frames: u32 = 60; - for _ in 0..frames { // a few frames to settle the title screen - while !bus.ppu.frame_done() { + for _ in 0..frames { + while !bus.frame_done() { let cycles = cpu.step(&mut bus); - bus.ppu.tick(&mut bus.cart, cycles as u64); // 3 PPU dots per CPU cycle + bus.tick_ppu(cycles as u64); } - if bus.ppu.take_nmi() { cpu.set_nmi(); } // loop bridges NMI - renderer.present(bus.ppu.framebuffer(), 256, 240); - bus.ppu.begin_frame(); + if bus.take_nmi() { cpu.set_nmi(); } + renderer.present(bus.framebuffer(), 256, 240); + bus.begin_frame(); } } diff --git a/src/ppu.rs b/src/ppu.rs index 31f0e9e..e7626d5 100644 --- a/src/ppu.rs +++ b/src/ppu.rs @@ -7,8 +7,22 @@ pub struct Ppu { status: u8, // $2002 oamaddr: u8, // $2003 oam: [u8; 256], // $2004 - scroll_x: u8, // $2005 - scroll_y: u8, + scanline: u32, + cycle: u32, + framebuffer: [u8; 256 * 240], + nmi_pending: bool, + frame_complete: bool, + odd_frame: bool, + t: u16, + x: u8, + pattern_lo_latch: u8, + pattern_hi_latch: u8, + bg_pattern_lo: u16, + bg_pattern_hi: u16, + bg_attr_lo: u16, + bg_attr_hi: u16, + tile_latch: u8, + attr_latch: u8, write_latch: bool, // shared W toggle for $2005/$2006, reset by $2002 read vram_addr: u16, // 14-bit data_buffer: u8, // $2007 read-buffer @@ -17,10 +31,68 @@ pub struct Ppu { } impl Ppu { + pub fn tick(&mut self, cart: &mut Cartridge, cycles: u64) { + for _ in 0..(cycles * 3) { + self.dot(cart); + } + } + + fn dot(&mut self, cart: &mut Cartridge) { + // vblank flag + NMI + if self.cycle == 0 && self.scanline == 241 { + self.status |= 0x80; + self.nmi_pending = true; + } + // clear vblank + if self.cycle == 0 && self.scanline == 261 { + self.status &= !0x80; + } + + let render_bg = self.mask & 0x08 != 0; + + if render_bg && self.scanline < 240 && self.cycle == 0 { + // fine-X scroll: consume the first `x` pixels of this scanline + self.bg_pattern_lo <<= self.x; + self.bg_pattern_hi <<= self.x; + self.bg_attr_lo <<= self.x; + self.bg_attr_hi <<= self.x; + } + + if render_bg && self.scanline < 240 && self.cycle < 256 { + self.render_bg_pixel(cart); + } + + if render_bg && (self.scanline < 240 || self.scanline == 261) { + self.bg_fetch(cart); + } + + // advance timing + self.cycle += 1; + let rendering = self.mask & 0x18 != 0; // bg or sprite render enabled + let line_len = if self.scanline == 261 && self.odd_frame && rendering { 340 } else { 341 }; + if self.cycle >= line_len { + self.cycle = 0; + self.scanline += 1; + if self.scanline >= 262 { + self.scanline = 0; + self.frame_complete = true; + self.odd_frame = !self.odd_frame; + } + } + } + pub fn new() -> Self { Self { ctrl: 0, mask: 0, status: 0, oamaddr: 0, - oam: [0; 256], scroll_x: 0, scroll_y: 0, + oam: [0; 256], + scanline: 0, cycle: 0, + framebuffer: [0; 256 * 240], + nmi_pending: false, frame_complete: false, odd_frame: false, + t: 0, x: 0, + pattern_lo_latch: 0, pattern_hi_latch: 0, + bg_pattern_lo: 0, bg_pattern_hi: 0, + bg_attr_lo: 0, bg_attr_hi: 0, + tile_latch: 0, attr_latch: 0, write_latch: false, vram_addr: 0, data_buffer: 0, vram: [0; 0x1000], palette: [0; 32], } @@ -62,16 +134,24 @@ impl Ppu { self.oamaddr = self.oamaddr.wrapping_add(1); } 0x2005 => { - if !self.write_latch { self.scroll_x = val; } else { self.scroll_y = val; } + if !self.write_latch { + self.t = (self.t & 0x7FE0) | ((val >> 3) as u16); // coarse X (bits 0-4) + self.x = val & 0x07; // fine X + } else { + self.t = (self.t & 0x0C1F) // keep fine Y + nametable Y + coarse X + | (((val & 0xF8) as u16) << 2) // coarse Y -> bits 5-9 + | ((val & 0x07) as u16) << 12; // fine Y -> bits 12-14 + } self.write_latch = !self.write_latch; } 0x2006 => { if !self.write_latch { - self.vram_addr = (self.vram_addr & 0x00FF) | (((val & 0x3F) as u16) << 8); + self.t = (self.t & 0x00FF) | (((val & 0x3F) as u16) << 8); // high 6 bits -> bits 8-13 } else { - self.vram_addr = (self.vram_addr & 0x7F00) | (val as u16); + self.t = (self.t & 0x7F00) | (val as u16); // low 8 bits + self.vram_addr = self.t; // v = t } - self.write_latch = !self.write_latch; + self.write_latch = !self.write_latch; } 0x2007 => { self.vram_write(cart, self.vram_addr, val); @@ -113,8 +193,8 @@ impl Ppu { fn nametable_offset(&self, cart: &Cartridge, addr: u16) -> usize { let a = addr & 0x2FFF; // fold the $3000-$3EFF mirror into $2000-$2FFF let table = match cart.mirroring() { - Mirroring::Horizontal => (a >> 10) & 1, - Mirroring::Vertical => (a >> 11) & 1, + Mirroring::Horizontal => (a >> 11) & 1, + Mirroring::Vertical => (a >> 10) & 1, Mirroring::OneScreenLow => 0, Mirroring::OneScreenHigh => 1, Mirroring::FourScreen => (a >> 10) & 3, @@ -129,6 +209,124 @@ impl Ppu { fn palette_read(&self, addr: u16) -> u8 { self.palette[palette_index(addr)] } + + fn bg_fetch(&mut self, cart: &mut Cartridge) { + // shift registers move left every dot + self.bg_pattern_lo <<= 1; + self.bg_pattern_hi <<= 1; + self.bg_attr_lo <<= 1; + self.bg_attr_hi <<= 1; + + if self.cycle == 255 && self.scanline < 240 { + self.copy_horizontal(); // dot 255: reload coarse X from t + } + + if self.cycle == 256 && (self.scanline < 240 || self.scanline == 261) { + self.increment_vertical(); // dot 256: advance to next scanline's tiles + } + + if self.cycle == 280 && self.scanline == 261 { + self.copy_vertical(); // dot 280 (pre-render): reload vertical bits from t + } + + match self.cycle % 8 { + 0 => self.tile_latch = self.nametable_read(cart, 0x2000 | (self.vram_addr & 0x0FFF)), + 2 => { + // attribute byte address (coarseY>>2, coarseX>>2, nametable select) + let addr = 0x23C0 + | (self.vram_addr & 0x0C00) + | ((self.vram_addr >> 4) & 0x38) + | (self.vram_addr & 0x07); + let byte = self.nametable_read(cart, addr); + // select the 2 bits for this tile's 2x2 block + let shift = ((self.vram_addr & 0x20) >> 3) | ((self.vram_addr & 0x01) << 1); + self.attr_latch = (byte >> shift) & 0x03; + } + 4 => { + let base = ((self.ctrl & 0x10) as u16) << 8; // BG pattern table select + let fine_y = (self.vram_addr >> 12) & 0x07; + self.pattern_lo_latch = cart.read_chr(base | ((self.tile_latch as u16) << 4) | fine_y); + } + 6 => { + let base = ((self.ctrl & 0x10) as u16) << 8; + let fine_y = (self.vram_addr >> 12) & 0x07; + self.pattern_hi_latch = cart.read_chr(base | ((self.tile_latch as u16) << 4) | fine_y + 8); + } + 7 => { + // reload shift registers with the fetched tile + self.bg_pattern_lo = (self.bg_pattern_lo & 0xFF00) | self.pattern_lo_latch as u16; + self.bg_pattern_hi = (self.bg_pattern_hi & 0xFF00) | self.pattern_hi_latch as u16; + // 0xFF/0x00 per attribute bit -> constant top bit during the tile's 8 pixels + self.bg_attr_lo = (self.bg_attr_lo & 0xFF00) | if self.attr_latch & 0x01 != 0 { 0xFF } else { 0x00 }; + self.bg_attr_hi = (self.bg_attr_hi & 0xFF00) | if self.attr_latch & 0x02 != 0 { 0xFF } else { 0x00 }; + self.increment_coarse_x(); + } + _ => {} + } + } + + fn increment_coarse_x(&mut self) { + if self.vram_addr & 0x001F == 0x001F { // coarse X wraps 31 -> 0 + self.vram_addr &= !0x001F; + self.vram_addr ^= 0x0400; // flip nametable X bit + } else { + self.vram_addr += 1; + } + } + + fn increment_vertical(&mut self) { + if self.vram_addr & 0x7000 != 0x7000 { + self.vram_addr += 0x1000; // fine Y++ + } else { + self.vram_addr &= !0x7000; // fine Y = 0 + let mut y = (self.vram_addr >> 5) & 0x1F; // coarse Y + if y == 29 { + y = 0; + self.vram_addr ^= 0x0800; // flip nametable Y + } else if y == 31 { + y = 0; + } else { + y += 1; + } + self.vram_addr = (self.vram_addr & !0x03E0) | (y << 5); + } + } + + fn copy_horizontal(&mut self) { // t's coarse X + nametable X -> v + self.vram_addr = (self.vram_addr & !0x041F) | (self.t & 0x041F); + } + + fn copy_vertical(&mut self) { // t's fine Y + coarse Y + nametable Y -> v + self.vram_addr = (self.vram_addr & !0x7BE0) | (self.t & 0x7BE0); + } + + fn render_bg_pixel(&mut self, _cart: &mut Cartridge) { + // fine-X offset: the first `x` dots render from the pre-fetched tile + let pattern_lo_bit = (self.bg_pattern_lo >> 15) & 1; + let pattern_hi_bit = (self.bg_pattern_hi >> 15) & 1; + let attr_lo_bit = (self.bg_attr_lo >> 15) & 1; + let attr_hi_bit = (self.bg_attr_hi >> 15) & 1; + + let pixel = (pattern_hi_bit << 1) | pattern_lo_bit; + let palette_bits = (attr_hi_bit << 1) | attr_lo_bit; + + let color = if pixel == 0 { + self.palette[0] // "universal background" transparent color + } else { + self.palette[(palette_bits * 4 + pixel) as usize] // background palette + }; + + let idx = (self.scanline as usize) * 256 + (self.cycle as usize); + self.framebuffer[idx] = color; + } + + pub fn frame_done(&self) -> bool { self.frame_complete } + pub fn take_nmi(&mut self) -> bool { let n = self.nmi_pending; self.nmi_pending = false; n } + pub fn framebuffer(&self) -> &[u8] { &self.framebuffer } + pub fn begin_frame(&mut self) { self.frame_complete = false; self.scanline = 0; self.cycle = 0; } + + + } fn palette_index(addr: u16) -> usize {