Industries · Fintech and trading

Fintech and trading

Nanosecond trading on FPGAs, certified secure chips in payment cards and HSMs, and single-function mining ASICs.

Lowest published STAC-T0 tick-to-trade I/O latency (2024)
13.9 ns minimum
EU clock rule for high-frequency traders
Within 100 µs of UTC, 1 µs stamps
First Bitcoin mining ASIC
January 2013, 130 nm
16 nm Bitcoin miners vs first 130 nm ASICs
~100× more energy efficient

At a glance

Where the flow bends

  1. 01Specification

    For trading hardware the spec is a time limit measured in billionths of a second. For payment chips it is a security certificate.

    Trading specs bound tick-to-trade latency and its variation. Secure-element specs name the certification target, such as a Common Criteria assurance level, EMVCo approval, or FIPS 140-3.

    Latency is specified as a distribution at a defined measurement point, for example from the last relevant inbound bit to the first outbound bit. For secure elements, the protection profile’s threats (leakage, probing, malfunction, abuse of functionality) become design requirements.

  2. 02Architecture

    Trading chips start acting on a message while it is still arriving, instead of waiting for all of it. Mining chips copy one small circuit thousands of times.

    Trading designs use cut-through processing: parse fields as bytes arrive, keep the order book in fast memory, and decide in a fixed number of cycles. The main platform choice is FPGA versus ASIC.

    The architecture avoids queues, arbitration, and shared resources on the critical path, because each adds variable delay. Off-chip memory lookups (for large books) and clock-domain crossings are the usual latency budget items.

  3. 03RTL design

    The logic is written as an assembly line where every step takes exactly the same time.

    RTL is deeply pipelined with no stalls, and protocol parsers for market data and order entry are written directly in hardware.

    Fixed-latency pipelines, parsing of several message types in parallel, and narrow datapaths that start work on early bytes. Mining RTL is a fully unrolled hash pipeline that finishes one hash per clock.

  4. 04Verification

    Tests run automatically every time the design changes, and they check the safety limits on trading as well as correctness.

    Teams use simulation with open-source tools such as Verilator and cocotb, plus continuous integration, and they verify pre-trade risk checks as carefully as functional behavior.

    Fast iteration matters more than methodology purity: some firms skip UVM for C++ and Python testbenches. Secure chips add a separate track of penetration testing and side-channel evaluation by accredited labs.

  5. 06Design for test

    On security chips, test features are also a door for attackers, so they are locked after the factory finishes testing.

    Scan chains and debug ports on secure elements must be disabled or protected after manufacturing test, because they can expose secrets.

    The security IC protection profile requires protection against abuse of functionality across the chip life cycle, which covers test modes. Test access is designed together with fuses, life-cycle states, and sensors.

  6. 10Clock tree synthesis

    Clocks are planned to avoid the small delays that happen when a signal passes between parts running on different clocks.

    Clock-domain crossings from the network receiver into the core logic add latency, so designs minimize them. Timestamp clocks are disciplined to UTC, often with PTP.

    Each asynchronous crossing costs synchronizer cycles and adds variation. Timestamping logic needs a clock traceable to UTC, and the regulation asks firms to identify exactly where in the system the timestamp is applied.

  7. 13GDS & tapeout

    Most trading hardware never becomes a custom chip at all. It is loaded onto reprogrammable chips that can change overnight.

    Trading logic usually ships as an FPGA bitstream. Custom ASIC tapeouts make sense only when logic is stable enough to justify mask costs. Bitcoin miners raced to each new process node.

    ASIC NRE and months of fab time conflict with strategy churn. In mining, the efficiency gain of each new node made the previous generation unprofitable, so tapeout timing decided which companies survived.

Finance uses custom chips for two very different jobs: speed and trust.

Speed. On an electronic exchange, the first order to arrive wins. Some trading firms put their trading logic in hardware sitting right on the network cable, so it can react to a price change in billionths of a second. Most of this runs on , chips that can be rewired by loading a new file.

Trust. The chip on your bank card holds secret keys. It has to keep them safe even if a thief takes the card to a lab and attacks it with equipment. These chips must pass independent security certification before banks can use them.

Trading firms are secretive about their hardware. This page sticks to what has been published: academic papers, public benchmarks, regulations, and what firms have chosen to say about themselves.

The key metric in trading is latency. Software on a CPU has latency on the order of tens of microseconds, and it varies from message to message. That is why FPGA cards are widely used to process market data.

Public benchmarks show how far hardware has pushed this. STAC-T0 measures from the last bit of inbound data needed for a decision to the first bit of the outbound order. In 2024 an FPGA solution posted a 13.9 ns minimum. That benchmark covers network I/O with trivial trading logic, so a real strategy adds its own cycles.

Consistency matters as much as raw speed. Hardware gives : the same message takes the same number of clock cycles every time.

Why FPGAs dominate trading. Exchange message formats change and strategies change faster. A 2014 FPGA trading paper states plainly that ASIC technology is not suitable for message decoding because the format of incoming messages often changes. Hudson River Trading describes trading as particularly time-to-market sensitive and uses FPGAs for its custom hardware.

When ASICs make sense. An trades and months of fab time for lower latency, power, and unit cost. That pays off for functions that change rarely, such as network interfaces or stable compute kernels. Startups have publicly pursued trading ASICs: in 2021 Bloomberg reported on one building an AI chip for high-frequency trading, with a foundry slated to manufacture it in 2022.

  • Being first. In trading, a few billionths of a second can decide who gets the deal.
  • Being consistent. A system that is usually fast but sometimes slow is risky, so designers remove anything that causes random delays.
  • Knowing the exact time. Regulators require trading firms to record exactly when things happened, so clocks must be accurate.
  • Resisting attackers. Payment and banking chips must keep keys secret even when attackers hold them in their hands.

Network integration. In tick-to-trade hardware, the Ethernet interface is part of the critical path. Low-latency results come from specialized network IP. A 2020 STAC-T0 result reached 24.2 ns minimum using a dedicated TCP core and 16-bit MAC/PCS cores on an FPGA.

Clocks and timestamps. In the EU, members of trading venues using high-frequency algorithmic trading must keep business clocks within 100 µs of UTC and timestamp with 1 µs granularity or better. The (IEEE 1588) can reach sub-microsecond synchronization, and sub-nanosecond with optimal network design.

Certified security. Payment card chips need an EMVCo security evaluation certificate for the IC, issued after testing by a recognized laboratory. Cryptographic modules such as are validated to FIPS 140-3 under the joint NIST and Canadian Cryptographic Module Validation Program.

Latency budget. A tick-to-trade path has SerDes and PCS/MAC receive, protocol parse, book lookup and update, decision, order encode, and transmit. Memory is often the bottleneck. One published FPGA design updates a book of 119,275 instruments in 253 ns on average using cuckoo hashing in 144 Mbit of off-chip QDR SRAM. Every clock-domain crossing and off-chip access shows up in the tail of the distribution.

Timestamp placement. RTS 25 requires firms to identify the exact point in the system where a timestamp is applied and show that it stays consistent. In hardware, that means timestamping at a fixed pipeline stage near the PHY.

Secure elements. The smart-card IC protection profile, written by Infineon, NXP, STMicroelectronics, and Inside Secure, claims EAL4 augmented by AVA_VAN.5 (advanced methodical vulnerability analysis) and ALC_DVS.2 (development security). Its functional requirements protect data against malfunction, leakage, physical manipulation, and probing, and prevent abuse of functionality. Leakage covers : power traces can reveal keys even when the crypto is a small fraction of total power. Certification is described in terms, and a ’s countermeasures live in RTL, layout, and analog sensors, so they cannot be added late.

  1. Trading: change often. Designs are loaded onto FPGAs and updated whenever the strategy or the exchange’s rules change.
  2. Trading: test constantly. Automated tests run on every change, and they check the built-in safety limits as well as the trading logic.
  3. Payment chips: invite attackers. Before approval, outside labs try to break the chip with real equipment.
  4. Mining chips: copy one block. The design is one small hashing circuit repeated across the chip.

FPGA flow. The front end matches the ASIC flow: RTL, simulation, and timing constraints. The back end is the FPGA vendor’s synthesis, place-and-route, and bitstream generation in place of floorplan, CTS, routing, signoff, and tapeout.

Verification for speed of change. Hudson River Trading describes co-simulation with cocotb, which drives simulator signals from Python, and Verilator, which turns Verilog into a C++ model, plus Jenkins to find regressions and try new random seeds. Verification checks that the device behaves as intended and that it obeys the firm’s risk checks.

Secure-element flow. The evaluation report for an EMVCo IC certificate includes vulnerability analysis and penetration testing by a recognized lab. Under Common Criteria, the development site’s own security is assessed too.

Timing closure for latency. Latency is pipeline depth × clock period, so teams close timing at a fixed, high clock and then fight for every stage removed. Cut-through designs can act on fields before the frame’s checksum arrives, so they need a way to cancel work on bad frames. Expect custom MAC/PCS IP, a minimal number of clock-domain crossings, and placement constraints that pin the critical path near the transceivers.

Methodology choices. HRT reports skipping UVM in favor of C++ and Python testbenches for code reuse and expressiveness. The trade is less standard methodology for faster turnaround.

Secure-element implementation. Masking, which splits secret data into randomized shares, is a standard DPA countermeasure. Implementation tools must not recombine those shares through logic optimization or sharing, so masked blocks often get special synthesis constraints. Test and debug access must be closed by life-cycle state after manufacturing test, since the protection profile requires preventing abuse of functionality. Resistance to malfunction and physical manipulation is part of the same requirement set. AVA_VAN.5 is decided by the lab’s attacks on real silicon.

Bitcoin mining is a race to guess a number. Miners compute a cryptographic fingerprint (SHA-256) of a block of data over and over, changing one number each time, until they find a result small enough to win. More guesses per second means more chances to win, and electricity is the main running cost.

Mining hardware moved from ordinary processors to graphics cards, then FPGAs, then custom chips. The first mining ASIC appeared in January 2013, and by mid-2015 miners had raced to 16 nm, the most advanced process of the day.

By 2017, the leading 16 nm miners were about 100 times more energy efficient than the first mining ASICs and about 8,000 times more efficient than graphics cards.

Mining is the clearest example of single-function silicon, and Michael Bedford Taylor documented its history in IEEE Computer.

  • One pipeline, copied. FPGA miners fully unrolled SHA-256 into 64 pipelined rounds, finishing one hash per clock cycle. Early ASICs mirrored that design.
  • ASICMiner (130 nm). Each USB-stick miner’s chip hashed at 330 MH/s at 1.05 V and 2.5 W, about 40× more energy efficient than a 28 nm GPU.
  • Butterfly Labs (65 nm). Each chip held 16 double SHA-256 pipelines on a 7.5 × 7.5 mm die. Taylor writes that preorder revenue presumably covered the $500,000 of mask NRE.
  • Bitmain Antminer S9 (16 nm). 189 ASICs in one machine delivered 13.5 TH/s at 1,323 W.

The mining story compresses ASIC economics into a few years.

  • Power estimation failures are fatal. Butterfly Labs’ chip drew four to eight times more power than expected. Every system had to be redesigned, and clearing the order backlog took nearly a year.
  • Time to market beats cost efficiency. HashFast and CoinTerra built cost-efficient 28 nm chips, but at more than 1.1 W per GH/s they were less energy efficient than BitFury’s 55 nm parts, which had shipped months earlier. That contributed to both companies going out of business.
  • Low voltage is the endgame. The 16 nm leaders run at ultralow voltages, and downward voltage scaling buys a few extra months of profitable life as mining difficulty rises.
  • Vertical integration. Leading miners co-design the ASIC, the machine, and the datacenter, which removes the need to support varied customer environments.

The same pattern appears in trading. When the function is fixed and the payoff per unit of latency or energy is clear, custom silicon wins. When the function keeps changing, the FPGA’s flexibility is worth more than the ASIC’s efficiency.

Sources

  1. Low Latency Book Handling in FPGA for High Frequency TradingMilan Dvořák, Jan Kořenek · IEEE DDECS 2014 (author copy, Brno University of Technology) · 2014FPGA cards widely used; software latency tens of µs and nondeterministic; ASICs unsuitable as message formats change; 253 ns book update.
  2. STAC Report: New STAC-T0 results with an Exegy/AMD FPGA solutionSTAC (Strategic Technology Analysis Center) · STAC Research · 2024Actionable tick-to-trade network I/O latency of 13.9 ns minimum on an FPGA.
  3. How We Verify Custom HardwareTodd Strader (Hudson River Trading) · HRT Beat · 2021FPGAs in trading; time-to-market sensitivity; cocotb, Verilator, CI; verifying risk checks.
  4. This Startup Is Building a Chip to Save Traders Vital MicrosecondsHooyeon Kim and Whanwoong Choi (Bloomberg) · Data Center Knowledge · 2021A startup developing an ASIC for high-frequency trading, with banks and quant firms in talks.
  5. STAC Report: New LDA/Xilinx solution under STAC-T0 (tick-to-trade network I/O)STAC (Strategic Technology Analysis Center) · STAC Research · 202024.2 ns minimum actionable latency using specialized TCP and 16-bit MAC/PCS IP cores on an FPGA.
  6. Commission Delegated Regulation (EU) 2017/574 (RTS 25) on the level of accuracy of business clocksEuropean Commission · legislation.gov.uk (The National Archives) · 2016High-frequency traders: 100 µs maximum divergence from UTC and 1 µs timestamp granularity; traceability rules in Article 4.
  7. IEEE 1588-2019: IEEE Standard for a Precision Clock Synchronization Protocol for Networked Measurement and Control SystemsIEEE Standards Association · IEEE · 2019PTP: sub-microsecond synchronization, sub-nanosecond under optimal network design.
  8. Chip & Platform Approval ProcessEMVCo · EMVCoSecurity evaluation certificates for payment ICs and platforms through recognized laboratories.
  9. Cryptographic Module Validation ProgramNIST Computer Security Resource Center · NISTCMVP validates cryptographic modules to FIPS 140-3 via accredited testing laboratories.
  10. Certification Report BSI-CC-PP-0084-2014: Security IC Platform Protection Profile with Augmentation PackagesFederal Office for Information Security (BSI) · Common Criteria Portal · 2014Smart-card IC protection profile: EAL4 augmented by AVA_VAN.5 and ALC_DVS.2; threats of malfunction, leakage, manipulation, probing.
  11. Introduction to differential power analysisPaul Kocher, Joshua Jaffe, Benjamin Jun, Pankaj Rohatgi · Journal of Cryptographic Engineering (open access; course copy, University of Michigan) · 2011Power measurements leak secret keys; attacks are practical and non-invasive; countermeasures.
  12. The Evolution of Bitcoin HardwareMichael Bedford Taylor · IEEE Computer (author copy, UC San Diego) · 2017CPU to GPU to FPGA to ASIC miners; unrolled SHA-256 pipelines; early ASIC case histories; efficiency by node.