Some of NVIDIA's CVDP Harnesses Pass Garbage

I ran NVIDIA’s CVDP harness for factorial_0001 against a module that outputs noise. It printed 100 failures, starting with Test 0 Failed: result = 14851409335729636181 do not matches reference 24, and then this:

Matched: 0 , Mismatched 100 - TEST FAILED.
** TESTS=1 PASS=1 FAIL=0 SKIP=0 **
1 passed

The harness logged every wrong answer and never failed the test. Four of the 39 problems I tried accept junk.

Setup

CVDP is NVIDIA’s benchmark for LLMs that write RTL. Each problem gives a model a spec, the model writes Verilog, and a cocotb harness simulates the result. The harness decides whether a model’s answer passes. I used the public v1.1.0 dataset without the agentic tasks, category cid003 (spec to RTL).

NVIDIA withholds the reference designs, so there is nothing to mutate. GateTruth audits benchmark testbenches by mutating the reference RTL and counting how many mutants get caught. It does this on RTLLM and says CVDP can’t be audited that way. So I skipped the reference and used junk.

For each problem I took the port list and parameters from an answer a model wrote and built a module with the same interface and no logic. Every output is independent of every input, in one of three patterns: all zeros, all ones, or an LFSR that changes every cycle. A harness that passes one of these can’t tell a working design from a broken one, whatever else it checks. I changed nothing in the harness except the file paths in its .env, which I pointed at a local directory. I ran it outside the official Docker image, with Icarus Verilog 13.0 and pytest, using cocotb 2.1.0 for the sweep. I reran the four problems below on cocotb 2.0.1 and pytest 8.3.2, the versions the dataset’s changelog lists, and every verdict was the same.

Results

I covered 39 of the 78 cid003 problems: the first 40 cid003 problems in the dataset file, minus ethernet_packet_parser_0001, whose interface I could not parse. That is 117 runs: 96 failed, 8 passed, 13 timed out. 28 problems rejected all three patterns. Seven are inconclusive because a junk design hung the harness until the timeout. Four accepted at least one:

Two bugs are behind those four.

factorial_0001 counts mismatches and prints the verdict:

if (mismatch==0):
  print(f"All {NUM_TESTS} test cases PASSED successfully.")
else:
  print(f"Matched: {match} , Mismatched {mismatch} - TEST FAILED.")

There is no assert or raise in the test or the runner. A cocotb test fails when it raises, and this one returns normally, so pytest reports 1 passed. Of the three junk patterns only the LFSR gets through: ones fails and zeros hangs until the timeout. The test picks its inputs with an unseeded random, so the reference value changes from run to run (I saw 24, 720 and 87178291200). The 100 failures don’t.

complex_multiplier_0001 and dbi_0001 have the same bug. Neither has an assert or raise in its test files or the runner. complex_multiplier_0001 ends with the same TEST FAILED print, and dbi_0001 prints [INFO] Test 'test_dbi_encoder' failed. and ends there. All three junk patterns pass both.

fibonacci_series_0001 is the other bug. It does have an assert, but the order is wrong:

# Check if overflow flag is set in DUT
if dut.overflow_flag.value == 1:
    dut._log.info("Overflow detected. Fibonacci sequence will reset.")
    break

# Assert to ensure DUT output matches the expected value
assert dut_output == a, f"Expected {a}, but got {dut_output} from DUT."

A design that holds overflow_flag at 1 leaves the loop as soon as it checks the flag, before the assert runs. With the ones module the log shows Overflow detected and both tests pass. Zeros and the LFSR fail.

Limits

I reported all four in NVlabs/cvdp_benchmark#60. #55 describes a test that “logs a failure without asserting” on a different set of problems.