A bridge connecting a schematic chiplet model to a detailed accelerator system
← Recent publications

Featured work · DAC 2026

ChiSER

Bridging the speed–fidelity gap in chiplet accelerator design.

A surrogate-enhanced, coarse-to-fine exploration framework for finding credible chiplet-based DNN accelerator designs without putting a slow reference model inside every search step.

Authors

Yi-Cheng Lo · Siva Satyendra Sahoo

Published at

63rd ACM/IEEE Design Automation Conference · DAC 2026

Presented

26–29 July 2026 · Long Beach, California, USA

01

The design gap

Why this work

Fast exploration and faithful evaluation rarely coexist.

Chiplet accelerator design spans partitioning, placement, dataflow, microarchitecture, interconnect bandwidth, and manufacturing cost. The search is enormous, and the model used to navigate it changes which designs appear optimal.

The speed–fidelity map used in the paper: ChiSER targets the high-speed region while recovering the intra-core detail that fast system estimators omit.
Fast system estimators

Broad search, blurred physics.

They evaluate candidates quickly, but can smooth over PE underutilization, partial-sum lifetimes, multicast traffic, and memory-port contention.

Fine-grain reference models

Credible detail, prohibitive runtime.

They capture intra-core behavior accurately, but minute-scale evaluations are too expensive for a large in-loop search.

Visual explanation of ChiSER bridging fast estimators and fine-grain reference models
ChiSER occupies the middle ground: reference-level intra-core fidelity at a speed suitable for design-space exploration.
MAP

Concept relationships

From fidelity gap to Pareto improvement

Trace how ChiSER moves faithful signals into the exploration loop.

The branches connect the modeling gap, the Intra-Core Surrogate, coarse-to-fine search, and the system-level improvements obtained when candidate ranking includes intra-core behavior.

ChiSER
Fast system estimatorsBroad search coverage with limited intra-core fidelity
Fine-grain reference modelsFaithful evaluation at impractical in-loop runtime
MisrankingCommunication and utilization errors distort the Pareto set
02

The approach

Coarse-to-fine exploration

Put faithful signals inside the search—not after it.

ChiSER introduces an Intra-Core Surrogate (ICS) alongside a fast system estimator. The surrogate supplies energy, delay, and feasibility signals while the search is still deciding how to partition and place work across chiplets.

ChiSER methodology connecting hardware enumeration, Gemini system estimates, the search algorithm, the intra-core surrogate, manufacturing cost, and Timeloop validation
Manufacturing cost, system traffic, and ICS estimates guide candidate ranking; Timeloop trains the surrogate and validates only the finalists.
03

Inside ICS

Architecture-aware surrogate

A compact model trained to notice what coarse heuristics miss.

Features

Hardware and workload structure

Array dimensions, buffer capacity, tensor sizes, kernel shapes, data-movement counters, and engineered ceil-fraction features expose divisibility and utilization effects.

Prediction

Hybrid XGBoost–MLP

Target-specific tree ensembles capture discrete boundaries; a lightweight neural head jointly predicts energy, delay, and mapping feasibility.

Deployment

Pruned, deterministic C++ runtime

Group-sparsity pruning removes redundant tree features before the surviving model is compiled into a low-overhead inference engine.

DAC presentation slide explaining ChiSER feature engineering, tree encoding, multi-target prediction, pruning, and deployment
Architecture-aware features feed tree-based encoders and a joint prediction head; structured pruning then produces a compact deterministic inference engine.
04

Measured impact

What changed

Local model fidelity reshapes the global Pareto frontier.

0.32 msper pruned surrogate query
0.997latency correlation with Timeloop
0.971energy correlation with Timeloop
62%lower MEDP for ResNet-50
44%lower MEDP for Transformer

MEDP is the manufacturing-weighted energy–delay product. The reported system-level reductions compare ChiSER-guided designs with designs selected using coarse heuristics.

Summary of the ChiSER problem, surrogate solution, and measured gains
Correcting intra-core estimates changes which chiplet designs survive the search and lowers the manufacturing-weighted energy–delay product.
05

Details

Model construction and validation

Follow ChiSER from feature engineering to Pareto-front refinement.

The sequence exposes the surrogate inputs, hybrid predictor, structured pruning, coarse-to-fine aggregation, and the final ResNet-50 and Transformer design-space results.

Presentation slide 1: ChiSER overview

01ChiSER overview

06

Watch

Chiplet exploration walkthroughs

See how intra-core fidelity changes the designs that survive.

Quick overview

ChiSER in under 90 seconds

The central problem, approach, and result in a short format.

Detailed presentation

High-fidelity chiplet exploration

Follow the motivation, method, surrogate design, and evaluation.

07

Publication

DAC 2026

ChiSER at the Design Automation Conference.

First page of the published ChiSER articleOpen published article ↗

Yi-Cheng Lo and Siva Satyendra Sahoo. “ChiSER: Surrogate-Enhanced Resource-aware Exploration of Chiplet-based Deep Neural Network Accelerators.” In the proceedings of the 63rd ACM/IEEE Design Automation Conference (DAC ’26), Long Beach, California, USA, 26–29 July 2026.

DOI · 10.1145/3770743.3803969 ↗

01

Yi-Cheng Lo

Graduate Institute of Electronics Engineering

National Taiwan University · Taipei, Taiwan

02

Siva Satyendra Sahoo

Pathfinding Co-optimization Technology & Systems

imec · Leuven, Belgium

Continue exploring the research landscape

AI for Design Automation →