From DSP feasibility to pre-silicon sign-off for a two-chiplet LiDAR ASIC
Six consecutive statements of work in fourteen months: algorithm-to-DSP feasibility, an emulation strategy study, software-driven pre-silicon verification of two chiplets, QEMU-to-RTL co-simulation, and DMA tooling the client later bought as source.
The challenge
A startup building a custom LiDAR control ASIC needed to know, in order: whether its spatial-aggregation algorithm would fit on a DSP, how to emulate the platform before silicon, whether two chiplets were correctly integrated, and how application developers could start before the parts arrived. It needed each answer before committing to the next.
What we owned
Architect. We transcoded the performance-critical stage of the client’s algorithm from Python to C, profiled it, and delivered a cycle-level DSP performance estimate with a written vectorization model, before any hardware decision was made.
Platform strategy. A study comparing a Synopsys Virtualizer virtual platform, Palladium-only emulation, and a SiliconScapes platform-evaluation approach, with licensing, throughput and fidelity trade-offs laid out. The client chose the hybrid.
Pre-silicon verification. For the ARM Cortex-M0 control chiplet, we converted UVM tests into processor-executed software tests to raise processor coverage, built a co-simulation framework to run software against gate-level peripheral models, and developed unit and system tests for DMA, UART, single-wire debug and the custom accelerators. For the DSP chiplet, we built a bus-functional model of the custom pipeline and a documented unit-test suite validating Tensilica TIE queue, APB and AXI integration under stall and back-pressure scenarios.
Co-simulation. Our host-speed processor framework developed the whole system test suite before silicon. QEMU was bridged to Xcelium over a SystemVerilog foreign-language interface for MMIO and interrupt-driven tests. A pre-silicon-arrival plan gave the client four fidelity-versus-speed options for early application development.
DMA tooling and validation. A DMA-330 microcode generator producing memory-image sequences for hardware-initiated transfers, an OCP SRAM bus-functional model on the client’s network-on-chip, and DMA functional and bandwidth validation across the pipeline. The utility was later purchased by the client as source.
Outcome
Six statements of work, each with written acceptance criteria and a joint go/no-go at the end, over fourteen months. Silicon issues found and fixed in RTL before tape-out. A tooling asset the client valued enough to buy outright.