SiliconScapesArchitecture · RTL · Prototype
Work / Tensilica DSP

Making Tensilica DSPs a platform: integration guidelines, core wrapping and accelerator attachment for a multi-DSP AR SoC

A platform-wide integration guideline every Tensilica core on the SoC follows, a repeatable core-wrapping procedure with foundry memory macros, and an accelerator attachment model built on TIE queues and shared data RAM.

ClientLow-power SoC for a small-form-factor AR device
Period2024 to 2025
Stages
  • Architect
  • Implement
  • Prototype
Tools
  • Cadence Xtensa XPG
  • TIE
  • Design Compiler
  • SpyGlass
  • TSMC N7 macros
  • SystemVerilog DPI
  • VCS

The challenge

The SoC carried several Tensilica DSPs with different configurations, plus custom hardware accelerators that had to attach to them. Each core arrived from the Cadence configuration flow as a “golden” deliverable with behavioral memories and its own interface conventions. The SoC team needed all of them to look like one family, to synthesize with real memory macros, and to be verifiable with software running on them.

What we owned

The platform guideline. One document defining how every DSP on the chip is integrated: reset sequencing with a software-controlled release bit, per-core clock gating, an alternate reset vector and a processor-ID register in a common control block so cores boot with unique identities before any software runs, synchronous AXI manager and subordinate interfaces with clock-enable strobes for integer clock ratios, iDMA trigger routing, and TIE queue conventions for accelerator attachment.

The core integration procedure. A written procedure and companion slide deck from Cadence XPG configuration through golden-core install, file import and renaming, replacement of memory placeholders with 7 nm foundry macros through an abstraction layer that preserves the Tensilica memory interfaces, subsystem instantiation, smoke tests, and co-simulation. The client’s own engineers now run it for new cores.

Physical memory organization. Data and instruction RAM bank structures against the foundry macro catalogue, with SECDED where the design required it.

Accelerator attachment. A hardware-accelerator proxy model covering AXI, TIE input and output queues, lookups and shared data RAM, verified with a hybrid flow: RTL where cycle behavior matters, transactional bus-functional models where functional behavior is enough, bridged through SystemVerilog DPI so C/C++ tests drive the RTL. This is the model our Cadence LIVE 2025 talk was built on.

Implementation support. Custom core generation through the Xtensa flow for hardware/software exploration, Design Compiler synthesis and SpyGlass lint for PPA on the integrated wrappers, and bring-up of a third-party NPU IP with the same memory-macro swap.

Outcome

A heterogeneous set of DSPs integrated to one standard, a procedure the client repeats without us, and an accelerator attachment approach that the Tensilica ecosystem has since heard about from a conference stage.