SiliconScapesArchitecture · RTL · Prototype
Work / XR silicon

Architecting a five-stage visual tracking accelerator for a multi-camera XR SoC

Hardware architect of record for a fixed-function accelerator that turns four concurrent camera streams into a real-time 6DOF tracking match set, coordinating the client's RTL designers across fifteen IP repositories.

ClientAR/XR headset OEM and its silicon design partner
Period2026, ongoing
Stages
  • Architect
  • Implement
  • Prototype
Tools
  • SystemVerilog
  • C++ golden model
  • draw.io
  • TSMC N7 macros
  • Git LFS
  • Xcelium

The challenge

Head tracking on a wearable has to run continuously at very low power. The client needed a fixed-function engine that would take pixels from a multi-camera cluster and deliver, every frame, a set of geometrically validated feature matches that the host fuses into a headpose update. The algorithms existed as reference software. The hardware did not, and the RTL team that would build it sat in another company.

What we owned

The specification. SiliconScapes is hardware architect of record. We authored and maintain the Hardware Architecture Specification, now past its eightieth page and revision 0.87, four block-level Micro-Architecture Specifications, an inter-stage interface specification, and the register map. Every change lands as a change note against a decision registry, so reviewers can see what moved and why.

The architecture. A five-stage streaming pipeline: keypoint detection, description, a single-pass sort-and-filter that selects a spatially uniform set without ever materializing a full frame of intermediate state, binned Hamming matching against banked on-chip descriptor storage, and RANSAC. Work is dispatched through a compact seven-operation instruction set from host-memory submission rings, with two independent stream contexts for real-time and best-effort traffic. A 252-bit rank-order descriptor packing roughly halves on-chip descriptor storage against the 512-bit source form, with a bit-exact codec.

The golden model. A bit-exact hardware model, held golden against the RTL on real image sequences and against the client’s own software reference, so every RTL release is judged on data rather than opinion.

The delivery method. The client’s designers implement the RTL across fifteen IP repositories. We run the workflow: tagged releases, pinned dependencies, staged design reviews, and integration benches at the seams between blocks. One such bench exposed a cross-IP addressing mismatch that neither block’s own testbench could see, because each modeled its own reading of the same sentence in the spec.

The prototype path. Our co-simulation platform stands in for the DMA subsystem, implementing the real register map, so firmware development runs against the architecture months before that RTL lands. Memory PPA inventories were produced against a 7 nm foundry macro catalogue.

Outcome

A controlled specification set that a distributed team builds from, a golden model that arbitrates every release, and a firmware platform that runs today. The engagement continues into RTL lockdown and FPGA prototyping.