Skip to content

RV32MFast: narrow-operand fast path for MUL (and DIV), disabled under data_ind_timing #2525

Description

@johnny7861532

Edited after re-checking our data: the first version overstated the timing / area x time results and some verification counts. The numbers below are the ones we stand behind.

Hi lowRISC team — I'm Johnny, founder of MorphusAI. We build AI systems for chip design,
and we've been evaluating our system on Ibex using your own flow. Before preparing a PR, we'd
like to ask whether you'd be interested in a small change to ibex_multdiv_fast.

Idea

  • In RV32MFast, MUL always spends 3 compute cycles (ALBL/ALBH/AHBL). When an operand's upper
    halfword is all zeros or all ones (i.e. it fits in 17 bits signed), the corresponding
    cross-product cycle can be skipped: for all zeros the cross product is zero, and for all ones
    its contribution to the low 32 bits is folded into the sign bit of the 17-bit multiplier input.
    MULH/MULHU/MULHSU are unchanged.
  • This follows the spirit of ⚡ Reduce latency of slow multiplier #590 (early termination in the slow multiplier, later gated by
    data_ind_timing) and the existing divide-by-zero early exit: the fast path is disabled when
    data_ind_timing_i is set, so data-independent-timing mode keeps today's cycle counts.
  • Separately, the divider can start at the first non-zero byte of |dividend| (bit 31/23/15/7)
    instead of always starting at bit 31, also only when data_ind_timing_i is clear and the divisor
    is non-zero.

Measurements (your syn/ scripts; Simple System CoreMark; ibex @ 46d0c75; each change compared with
the unmodified config in the same flow)

  • small + MUL fast path: CoreMark cycles -4.6% (4,063,528 -> 3,875,608, cycle-exact).
    Area +0.5% to +1.3% across 5 syntheses with logically equivalent rewrites elsewhere in the RTL.
  • maxperf: the MUL change has no effect (single-cycle multiplier; identical synthesized netlist).
  • DIV change: 0 CoreMark cycles in every config we ran. On riscv-tests qsort (one static DIV,
    executed 293 times; no MUL), the MUL + DIV build saves 4,808 of 190,146 cycles (-2.5%) versus
    the unmodified build, exactly matching a prediction from the instruction trace. Since qsort
    executes no MUL we attribute this to the DIV change (not measured separately). That run used a
    custom RV32MFast + writeback build rather than an official config.
  • Timing: we could not resolve a min-period effect in either direction. In this flow the period
    of the same design varies by about 13% (unmodified small and maxperf) and by up to about 16%
    (patched trees) across logically equivalent rewrites, so we are not claiming any frequency
    change.

Verification so far

  • Module level: 498,048 directed and random MUL/MULH/DIV/REM transactions for the proposed changes
    (MUL, DIV, both), for both RV32MFast and RV32MSingleCycle elaborations, against an independent
    C++ golden model (signed/unsigned corners, overflow, divide by zero, backpressure, enable stalls,
    async reset mid-operation): 0 mismatches. Fault injection (flipped result bit, removed DIT
    guard) is caught.
  • With data_ind_timing_i set, per-transaction latency is identical to the unmodified RTL.
  • Your CI on small and maxperf for each change individually (Verilator/Verible lint,
    riscv-compliance, Spike CoreMark cosim, CSR testbench, lint of the other configs): same results
    as the unmodified RTL. In our environment rv32i compliance shows the same 4/48 failures with and
    without the change. The combined MUL + DIV build has not been through CI yet.
  • Both patches apply cleanly to current master (af456e2). Module-level Verilator 5.052 -Wall lint
    of ibex_multdiv_fast there (RV32MFast and RV32MSingleCycle) shows no new warnings versus
    unmodified master; the warnings that do appear are identical before and after (UNUSEDPARAM from
    ibex_pkg in a standalone lint, plus PINMISSING from our test wrapper). Full-core CI on master
    has not been run yet.
  • Not yet run: your UVM / riscv-dv regression, or commercial synthesis.

Questions

  1. Would you accept a variable-latency MUL fast path in RV32MFast if it is disabled under
    data_ind_timing_i?
  2. Would you prefer MUL and DIV as separate PRs?
  3. What additional verification would you want to see?

Disclosure: these changes were produced with MorphusAI's AI chip design system and reviewed by me;
I take responsibility for them. Happy to share the patches and the measurement data.

My Environment

EDA tool and version:

  • Synthesis/timing: your syn/ scripts with sv2v 6662fa5, Yosys 0.62 and OpenSTA 3.1.0 (built from
    OpenROAD's pinned submodule 43177bba), Nangate45. We could not evaluate the pinned Nix shell, so
    these tool versions differ from it.
  • Simulation/CI: Verilator 4.210, lowRISC GCC 10.2.0 (20220210-1), FuseSoC 2.4.3, Edalize 0.6.8,
    Verible v0.0-3622, Spike cosim as in your CI.
  • Module-level tests: Verilator 5.052.

Operating system:
Ubuntu 24.04.5 LTS (cloud VMs) for the flows above; macOS for local module tests.

Version of the Ibex source code:
46d0c75 for all measurements. The changes touch only
rtl/ibex_multdiv_fast.sv; both patches also apply cleanly to af456e2.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions