Industries · Automotive

Automotive

Chips that must detect their own faults, survive under the hood and ship with near-zero defects.

AEC-Q100 Grade 0 ambient range
−40 °C to +150 °C
ASIL D single-point fault metric target
≥ 99%
AEC-Q100 HTOL sample
77 parts × 3 lots, 0 fails
Tesla FSD chip (2019)
6 billion transistors, 260 mm²

At a glance

Where the flow bends

  1. 01Specification

    The spec says how dangerous a failure would be, how hot the chip will get, and how the chip must protect itself from hackers.

    Add safety goals with an ASIL (A–D) from the hazard analysis, an AEC-Q100 temperature grade, a mission profile for years of use, target hardware metrics, and cybersecurity goals from a threat analysis.

    Derive chip-level safety requirements from the item’s safety goals: target SPFM, LFM and PMHF for the ASIL, fault-tolerant time intervals, safe states, and assumptions of use if the SoC is a safety element out of context. Add ISO/SAE 21434 cybersecurity goals from the TARA.

  2. 02Architecture

    The chip gets built-in watchdogs: doubled processors that check each other, memory that fixes its own errors, and a timer that notices if software freezes.

    Add safety mechanisms: dual-core lockstep CPUs, ECC on memories, watchdogs, and often a separate safety island that supervises the big compute blocks.

    Partition ASIL-D control onto lockstep cores and a safety island with independent clock and power; decide diagnostic coverage per block; plan freedom from interference between mixed-criticality software and hardware partitions.

  3. 03RTL design

    Engineers write the checking circuits right alongside the normal ones.

    RTL includes ECC encoders and checkers, comparators for lockstep, error signaling to a central fault collector, and hooks for self-test.

    Every safety mechanism needs an error output that reaches a fault-handling path, plus a way to test the mechanism itself, since an untested checker is a latent fault.

  4. 04Verification

    Engineers deliberately break the design in simulation thousands of times to prove the checking circuits catch the problems.

    Fault-injection campaigns insert stuck-at and transient faults and classify each one as safe, detected or dangerous. The results feed the safety metrics.

    Run simulation-based fault campaigns per safety mechanism, use formal to prove unobservable faults safe, and roll the classified results into the FMEDA to compute SPFM and LFM against the ASIL targets.

  5. 06Design for test

    Test circuits are built in so the factory can catch nearly every bad chip, and so the chip can test itself later in the car.

    Very high scan coverage (stuck-at and transition faults) supports near-zero defects; logic and memory BIST let the chip test itself in the field to find latent faults.

    Fault coverage targets are higher than consumer parts, with at-speed patterns, outlier screening (Part Average Testing) and in-system LBIST/MBIST whose run time must fit the diagnostic test interval and power budget.

  6. 12Signoff

    Timing and reliability are checked at higher temperatures and over more years than for a phone chip.

    Corners extend to the temperature grade (up to 150 °C ambient for Grade 0) and the mission profile sets electromigration and aging margins.

    Sign off EM, IR and aging against the automotive mission profile and confirm safety mechanisms meet timing in every corner, since a missed check is a dangerous fault.

  7. 13GDS & tapeout

    After manufacturing, chips go through months of stress testing before a carmaker will use them.

    Parts are qualified to AEC-Q100, including 1,000-hour high-temperature operating life on samples from three lots with zero failures allowed.

    Plan AEC-Q100 test groups (environmental, lifetime, package, electrical verification), zero-defect practices from AEC-Q004 in production, and a safety case with FMEDA and fault-injection evidence for the customer’s assessment.

A chip in a car may help steer, brake or drive. If it fails at the wrong moment, people can get hurt. So car chips are designed to notice their own faults and react safely, to survive years of heat, cold and vibration, and to almost never leave the factory broken.

The rulebook for the first part is called . Each possible danger gets an rating from A to D, based on how bad the harm would be, how often the situation happens, and whether a driver could still control the car. The higher the rating, the more checking the chip needs.

ISO 26262 is the functional safety standard for road vehicles. Its 2018 second edition widened the scope to all road vehicles except mopeds and added Part 11, guidelines on applying the standard to semiconductors. The hazard analysis rates each hazard for severity (S), exposure (E) and controllability (C), and a table combines them into an ASIL for each safety goal.

For chip hardware, the ASIL sets numeric targets:

MetricASIL BASIL CASIL D
≥ 90%≥ 97%≥ 99%
≥ 60%≥ 80%≥ 90%
under 100 FITunder 100 FITunder 10 FIT

Separately, every automotive IC is stress-qualified to .

The chip team’s safety deliverables are with quantified diagnostic coverage, an that rolls failure rates and coverage into SPFM, LFM and PMHF, and fault-injection evidence that the coverage claims hold. An analysis of RISC-V automotive safety argues that the dominant costs are engineering activities: FMEDA generation, diagnostic coverage analysis, tool qualification, safety-case documentation and fault-injection campaigns.

ISO 26262 covers malfunction. Two other standards sit beside it: ISO 21448 (SOTIF) addresses perception and decision failures outside the random-hardware-fault model, which matters for autonomy, and ISO/SAE 21434, released in August 2021, sets cybersecurity engineering and risk analysis principles.

  • Safety: the chip must catch its own faults.
  • Harsh conditions: under the hood, a chip may need to work from −40 °C to +150 °C.
  • Zero defects: carmakers count bad chips per million shipped, and the goal is none.
  • Hackers: connected cars need chips that resist attack.
  • Huge computers: driver-assistance and self-driving systems need the power of a data-center chip with the reliability of a brake controller.

Temperature grades. AEC-Q100 defines four ambient ranges: Grade 0, −40 °C to +150 °C; Grade 1, −40 °C to +125 °C; Grade 2, −40 °C to +105 °C; Grade 3, −40 °C to +85 °C. High-temperature operating life (HTOL) runs 1,000 hours at the grade’s maximum ambient on 77 parts from each of three lots, with zero failures allowed.

Zero defects. The AEC’s zero-defects framework is a menu of practices across process design, product design, production and improvement. For designers, it calls for design for test that reaches “as many nodes as possible,” including scan stuck-at and transition fault coverage, and built-in self-test. In production, Part Average Testing removes statistical outliers from the population before they ship. The target is as close to zero as possible.

Cybersecurity. ISO/SAE 21434 sets high-level principles for threat analysis and risk assessment () but does not prescribe use cases or automation. At chip level, TARA outputs become requirements for secure boot, key storage and debug access.

Autonomy compute. High-performance multicores and GPU-class accelerators are increasingly the only way to reach the performance autonomous cars need, but the safety support they offer is uneven. A complements them with watchdogs, monitoring, test orchestration and diverse redundant execution.

Lockstep. ASIL-D components are generally deployed on dual-core CPUs: the same software runs on two identical cores with a time stagger, so a fault hitting both at once produces different errors that a comparator catches. Watchdogs should be independent of what they monitor, with their own clock and power supply where possible.

Safety goals and temperature grades go into the spec. The architecture adds checkers: twin processors, self-correcting memory and watchdog timers. In verification, engineers break the design on purpose thousands of times in simulation to show the checkers catch problems. Extra test circuits let the factory catch nearly every defect, and let the chip test itself later in the car. Finally the chips spend months in stress tests before approval.

StageAutomotive addition
SpecSafety goals and ASIL, temperature grade, mission profile, security goals
ArchitectureLockstep cores, ECC, watchdogs, safety island
RTLSafety mechanisms with error outputs to a fault collector
VerificationFault-injection campaigns feeding SPFM and LFM
DFTHigh scan coverage; in-field logic and memory BIST
SignoffCorners to the temperature grade; mission-profile EM and aging
TapeoutAEC-Q100 qualification; safety case

ISO 26262 highly recommends during IC development to evaluate how well the design handles random hardware failures. A campaign inserts faults into ports, flip-flops and wires, then classifies each by whether it reaches an output and whether a safety mechanism flags it.

Fault classification. Faults observed at the output but not at the error signal are dangerous undetected faults. Faults seen at neither are potential latent faults that need another diagnostic, such as . Simulation-based campaigns are not exhaustive, so unobserved faults traditionally need manual analysis; formal tools can prove many of them untestable by design. In one published campaign on a SECDED ECC block, 512 of 1,726 faults corrupted output data without raising an ECC error.

Campaign cost. Transient-fault campaigns grow with design size and simulation length. Fault-list pruning helps; dynamic HDL slicing cut injections by up to 10% on an industrial core.

DFT doubles as safety. The same scan and BIST infrastructure that drives DPPM down at the factory provides in-field diagnostics. BIST trades area, test time and supply requirements against fault coverage and lower test cost. In the field, LBIST scrambles functional state, so it runs in windows when the function can be paused, and the run must fit the safety concept’s diagnostic interval.

In 2019 Tesla described its own self-driving chip at the Hot Chips conference. Each car computer carries two of these chips, each with its own power supply, so one can back up the other. The chip is qualified to the AEC-Q100 car-chip stress tests, and Tesla says it took 14 months from architecture to tapeout.

Most of the chip uses proven, licensed building blocks. The part Tesla designed itself is a neural-network engine that does the heavy math for recognizing what the cameras see.

Published specs:

  • 14 nm FinFET CMOS, 260 mm², 6 billion transistors, 37.5 × 37.5 mm FCBGA package, AEC-Q100.
  • 12 CPU cores, a GPU, two neural network accelerators (NNAs), an image signal processor and an H.265 encoder, plus safety and security blocks.
  • Each NNA: a 96 × 96 multiply-accumulate array at 2 GHz+, 36.8 TOPS, and 32 MB of SRAM.
  • Goals: over 50 TOPS, about 80% utilization, under 40 W per chip, and under 100 W for the computer.

The computer board has dual redundant SoCs, redundant power supplies, and overlapping camera fields with redundant paths.

Two design choices stand out. First, redundancy is at system level: Tesla listed “lower part costs to enable redundancy architectures” as a platform goal and designed the chip to be “modular to enable various platform redundancy uses.” This matches the broader pattern for autonomy, where high-performance compute is paired with separate safety supervision because the safety support built into such compute is uneven.

Second, schedule risk was cut by limiting custom work. Standard functions (CPUs, GPU, ISP, video encoder, memory controller, PHYs, interconnect) came from proven IP, and the team chose simpler clock and power distribution. The custom NNA uses a single clock domain, state-machine control and DVFS-enabled power and clock distribution, and keeps programs resident in SRAM to avoid DRAM traffic.

What the slides do not include is equally typical: no ASIL rating, FMEDA or fault-injection results are published. Automotive safety cases go to customers and assessors, so public case studies show the architecture and keep the safety evidence private.

Sources

  1. Assessment of Safety Standards for Automotive Electronic Control Systems (DOT HS 812 285)Qi D. Van Eikema Hommes · National Highway Traffic Safety Administration · 2016How ISO 26262 assigns ASIL from severity, exposure and controllability.
  2. ISO 26262WikipediaEditions and scope; Part 11 guidelines on semiconductors in the 2018 edition.
  3. Formal Assisted Fault Campaign for ISO26262 CertificationNitin Ahuja, Mayank Agarwal and Sandeep Jana · DVCon Europe proceedings · 2019SPFM/LFM/PMHF targets by ASIL; fault classification; ECC case.
  4. Accelerating Transient Fault Injection Campaigns by using Dynamic HDL SlicingAhmet Cagri Bagbaba, Maksim Jenihhin, Jaan Raik and Christian Sauer · arXiv · 2020ISO 26262 recommends fault injection; campaign cost.
  5. AEC-Q100 Rev-J1: Failure Mechanism Based Stress Test Qualification for Integrated Circuits in Automotive ApplicationsComponent Technical Committee · Automotive Electronics Council · 2026Temperature grades; HTOL conditions and sample sizes.
  6. AEC-Q004: Automotive Zero Defects FrameworkComponent Technical Committee · Automotive Electronics Council · 2020Zero-defect practices: BIST, DfT coverage, Part Average Testing.
  7. Envisioning a Safety Island to Enable HPC Devices in Safety-Critical DomainsJaume Abella, Francisco J. Cazorla, Sergi Alcaide, Michael Paulitsch, Yang Peng and Inês Pinto Gouveia · arXiv · 2023HPC devices in cars; dual-core lockstep for ASIL D; watchdogs; safety islands.
  8. Security Risk Analysis Methodologies for Automotive SystemsMohamed Abouelnaga and Christine Jakobs · arXiv · 2023ISO/SAE 21434 release and its TARA principles.
  9. Compute and Redundancy Solution for the Full Self-Driving Computer (Hot Chips 31 slides)Pete Bannon, Ganesh Venkataramanan, Debjit Das Sarma, Emil Talpes, Bill McGee and team (Tesla) · Hot Chips 31 · 2019FSD chip and computer: redundancy, AEC-Q100, NNA design, schedule.
  10. RISC-V Functional Safety for Autonomous Automotive Systems: An Analytical Framework and Research Roadmap for ML-Assisted CertificationNick Andreasyan, Mikhail Struve, Alexey Popov, Maksim Nikolaev and Vadim Vashkelis · arXiv · 2026Certification cost drivers; SOTIF alongside ISO 26262.