Memora8 FPGA Implementation Overview: Eight-Core Artix-7 Processor, Memory, and Verification Status

AI Summary: As of October 4, 2026, the Memora8 implementation targets an AMD Artix-7 XC7A200T on the BX72 board and contains a cluster of eight 32-bit processor cores. The design includes 1 MiB of banked on-chip SRAM, eight MMU page windows per CPU, SYSTEM BUS devices, PageMover, a DDR3/MIG connection, UDP debug, and TTY output collection. The completed routed build passed the reported setup, hold, pulse-width, routing, DRC, and bus-skew checks at 125 MHz, with small positive timing margins. The UDP integration suite passed against the Verilator design, but it bypasses the physical Ethernet MAC/RGMII and models DDR3/MIG; physical board-level Ethernet and DDR operation, execution on the FPGA, and cryptographic accelerators remain unverified or planned. No claim of complete formal verification is made.

Memora8 has progressed from an instruction-set and simulator project to an FPGA implementation with an eight-core processor cluster, banked memory, system devices, and a host-side debug path. This overview records what is implemented, what has passed simulation and implementation checks, and what still requires board-level validation.

The status described here corresponds to implementation commit 677e631 and the completed FPGA build verified on October 4, 2026. Planned devices are identified separately from the current hardware. For earlier context, see the reports on bringing Memora8 from simulator to Artix-7 hardware and the K4,4 memory topology and 125 MHz timing results.

Eight-Core 32-Bit Processor Cluster

The current implementation contains one cluster of eight 32-bit processor cores, CPU0 through CPU7, targeting an AMD Artix-7 XC7A200T FPGA on the BX72 board. Each core has its own pipeline, architectural register file, eight MMU page windows, and trap-handler entry.

The instruction set includes integer arithmetic, comparisons, shifts, branches, word loads and stores, multiplication, division, remainder, and SYSTEM BUS operations. Each CPU has 32 architectural 32-bit registers. Multiplication and division use multicycle execution units.

The implementation also includes registered instruction-fetch routing, MMU-context validation, and separate MEM operand control for all eight CPUs. Simulation exercises dependent load, arithmetic, multiply, divide, and store sequences throughout the cluster.

A redirect to address zero places a core into a sleep state. SYSTEM BUS control can wake or restart it. When a core is stopped, accepted work is drained before completion is reported. The Memora8 ISA documentation describes the architectural instruction set; this implementation report focuses on the current processor and FPGA integration.

Memory Map and Page Model

The FPGA build has 1 MiB of on-chip SRAM organized as 16 banks, with eight pages of 8 KiB per bank. That is 128 physical pages, numbered 0 through 127. Each page contains 2,048 32-bit words and has one memory port.

Resource Current configuration
On-chip SRAM 16 banks × 8 pages × 8 KiB = 1 MiB
Physical SRAM pages 128 pages, numbered 0–127
Page organization 2,048 × 32-bit words; one memory port per page
CPU address space 32-bit virtual addresses, mapped through eight windows per CPU
MMU window size 8 KiB; up to 64 KiB mapped per CPU at a time
External DDR3 1 GiB, 32-bit interface; PM1 addresses 131,072 pages of 8 KiB

SRAM uses a fixed shared-bank topology. Each CPU is physically connected to four banks, and each bank connects a fixed pair of CPUs. Mapping checks both connectivity and page occupancy. A page has one active user at a time; page storage itself does not contain an owner-selection controller. Different pages can still be accessed concurrently.

When instruction fetch and data access target the same page, instruction fetch has priority and the data access waits. Accepted writes are preserved. Once a read request is admitted, the BRAM read takes one cycle; this describes the memory access latency, not one-cycle instruction execution.

SRAM contents must be loaded before execution. CPUs do not access external DDR3 directly through ordinary loads and stores. PageMover provides the transfer path:

  • PM0 copies SRAM pages.
  • PM1 transfers pages between SRAM and DDR3 through the Memory Interface Generator (MIG).

Both use the shared 32-bit PageMover data path and are serialized in the current design. The separate Memora8 memory architecture report discusses the bank topology and timing experiment in more detail.

SYSTEM BUS and Device Map

Each CPU has a 32-word SYSTEM BUS buffer. Words 0–30 carry data. Writing word 31 with (CMD << 10) | DEVICE launches an operation. A device address contains a five-bit group and a five-bit device index. Ordinary responses put status in DATA0 and results beginning at DATA1; SB_ECHO preserves the data buffer.

Group Implemented functions
0 Identification and available-group mask; ticks; system and cluster atomic counters; two cluster timers; CPU sleep, wake, and state; exchange registers; MMU context and page-window control; MTR entry; restart; buffer clear; EVENT_POOL; SB_ECHO
1 Read-only PAGE_OCCUPANCY bitmap for the 128 SRAM pages
6 Cluster register block with 32 × 32-bit registers
7 Exported register block with four 32-bit registers per CPU
8 Two atomic FIFOs, each with 4 KiB storage and configurable item width
9 TTY_TX, device 0: 64 × 16-bit FIFO
16 PM0 at 0x200; PM1 at 0x201
17 Clock device

CPU-local facilities also include two atomic counters and one timer per CPU. The exchange-register bank retains 32 words per CPU and is separate from the four-register exported block. EVENT_POOL reports pending device events.

Group 2 and temporary register blocks in groups 3–5 are absent. Full device-passport reporting and parts of the timer/clock ABI still need integration. The SYSTEM BUS and SB/SBI reference provides the interface-level documentation.

UDP Debug and Board Peripherals

UDP debug is an independent SYSTEM BUS client with its own buffer. It supports page upload and download, CPU run/stop/reset, MMU control, PageMover operations, EVENT_POOL inspection, and generic bus commands. The FPGA network top does not expose direct CPU debug ports.

TTY output accumulates in a 256-byte debug buffer and is returned when the client requests it. The FPGA design includes the BX72 Ethernet transport for UDP debug and the DDR3 MIG connection. It does not yet include a general-purpose operating-system network driver or network stack.

The following peripherals are not implemented in the current build:

  • SD or RAW storage;
  • LVDS/HDMI display;
  • laptop keyboard and touchpad support.

The present implementation comprises the processor cluster, SRAM, MMU, PageMover, SYSTEM BUS devices, UDP debug transport, and TTY collection described above. Planned and deferred features are not part of this completed build.

Cryptographic Accelerators: Planned, Not Implemented

The planned Memora8 and Reganta cryptographic profile consists of four accelerators. None is implemented in the current FPGA build.

Device Purpose Control and results Bulk data path
BLAKE3-256 Hash pages, modules, and other data; derive keys according to the cryptographic profile SYSTEM BUS PageMover stream or successive SYSTEM BUS chunks
Ed25519 Sign and verify firmware, system modules, updates, and credentials SYSTEM BUS only No PageMover access; operates on a small canonical object containing a hash
X25519 Establish a shared secret for session keys SYSTEM BUS only No PageMover access
XChaCha20-Poly1305 Authenticated encryption and decryption SYSTEM BUS Streaming input and output through PageMover

The proposed devices use the existing SYSTEM BUS command word and data buffer, without direct CPU channels or additional SRAM-bank ports. Device addresses and command numbers have not yet been assigned.

For BLAKE3, a SYSTEM BUS command would start an operation, PageMover would supply bulk data, and SYSTEM BUS would return the 256-bit digest as eight words. Small inputs could instead be supplied as successive chunks while the accelerator retains its hash state. The proposed sequence is BEGIN, UPDATE, and FINAL. With one word reserved for chunk length, an UPDATE can carry up to 120 bytes; a 256-byte input would use chunks of 120, 120, and 16 bytes.

Large objects would be hashed with BLAKE3 before signing. Ed25519 would sign a canonical object containing a domain identifier, format version, original data length, and BLAKE3 digest. Verification would recompute the digest and check the signature over the same encoded object. This design uses ordinary Ed25519, including its internal SHA-512 operation. The external BLAKE3 digest does not replace that internal hash and does not make the operation Ed25519ph. Ed25519 uses a 32-byte public key, a 32-byte private seed, and a 64-byte signature; verification does not require a private key.

X25519 would receive its parameters through SYSTEM BUS and return a 32-byte shared secret. Session-key derivation would incorporate connection context using BLAKE3 derive_key. XChaCha20-Poly1305 uses a 32-byte key, a 24-byte nonce, and a 16-byte authentication tag. SYSTEM BUS would carry control and associated data, while PageMover would carry the message stream. A nonce must not repeat under the same key, and decrypted data must remain unavailable to consumers until its authentication tag has been verified.

Before implementation, the design still needs exact signature encoding, secure key storage, a randomness source, secret-lifetime rules, and authenticated-plaintext release semantics. The existing BLAKE3_PAGE_VERIFY contract also needs to be reconciled with the proposed streaming interface. The plan excludes MACC8/DOT8 devices and dedicated arithmetic channels attached to SRAM pages.

Validation and Timing Results

The Verilator executable runs the current SystemVerilog processor and device logic; it is not a separate software implementation of the instruction set. The UDP integration suite passed on October 4, 2026, reporting MR8_UDP_TEST_OK.

The suite covers page transfers, PM0, PM1 with delayed DDR responses, MMU and page occupancy, EVENT_POOL, TTY, execution on all eight CPUs, packed SB/SBI commands, AUIPC, sleep/wake, SB_ECHO, and client sequence persistence. Earlier runs of language and language_lib completed successfully without recorded traps.

The completed Vivado implementation for the XC7A200T-2 passed these checks:

Check Result
Setup timing WNS +0.049 ns; TNS 0; no failing endpoints
Hold timing WHS +0.039 ns; THS 0; no failing endpoints
Pulse width No failing endpoints
Routing All 156,289 routable nets fully routed; zero routing errors
DRC Zero errors; warnings remain in the report
Bus skew All eight constraints met; worst skew 0.770 ns against an 8.000 ns limit; minimum slack +7.230 ns
Bitstream generation Completed successfully; output mr8/build/bit/mr8.bit

The routed build meets timing at a 125 MHz system clock. The positive setup and hold margins are small, so timing must be rechecked after implementation or constraint changes. These results establish timing closure for this routed build under its applied constraints; they are not a guarantee that future revisions will meet timing.

The functional results have important limits. The UDP test harness bypasses the physical Ethernet MAC/RGMII interface and replaces DDR3/MIG with a memory model. Physical DDR operation, Ethernet electrical timing, and execution on the FPGA still require board-level validation. The Vivado flow produces synthesis, placement, routing, utilization, and DRC reports, rejects negative setup or hold slack before bitstream generation, and checks bus skew separately for this build.

These are test and implementation results, not exhaustive correctness claims. No claim of complete formal verification is made. The Memora8 FPGA bring-up report describes the earlier single-CPU board profile; this overview records the current eight-core implementation and its remaining validation work.

Current Status in Brief

Memora8 now has an eight-core FPGA implementation, a 1 MiB banked SRAM system, MMU page windows, SYSTEM BUS control and data devices, PageMover connectivity to DDR3/MIG, a UDP debug client, and a routed build that closes timing at 125 MHz. Simulation exercises the cluster and a broad set of device and memory operations.

At the same time, the verification boundary remains explicit: the reported UDP suite substitutes models for physical DDR and the Ethernet MAC/RGMII path, and the cryptographic devices remain planned. Board-level operation and further validation are still required. Keeping implemented features, simulated behavior, physical tests, and roadmap items separate is essential to an accurate account of the Memora8 system.