falcon_mdf Field Reference
Capability reference · main

falcon_mdf, end to end

Every capability the crate and its viewer ship today, what is genuinely missing, and how both compare against asammdf, mdfreader and the Rust MDF field — each row carrying the file, test or measurement it rests on. This revision covers dev/0.7.0: all six zip types now write, the library reads from any byte source and the browser streams files and URLs larger than its memory, exports stopped refusing variable-length arrays and MAT text, both viewers list Ethernet and FlexRay, and an independent C++ writer, mdflib, joined the test set and found three FlexRay bugs.

At a glance

falcon_mdf reads, writes, decodes and inspects ASAM MDF measurement files in Rust, and ships a desktop viewer, a Python binding and a browser viewer alongside the library. It is designed around one promise: a channel decodes to the right values, or reading it fails with a reason that names the feature.

Published: 0.6.0 on crates.io; 0.7.0 ready on a branch

falcon_mdf 0.6.0 is on crates.io and docs.rs, with viewer binaries on the GitHub release. Everything this page adds since is on dev/0.7.0, verified and not yet published.

Where it leads

Capabilities that no comparable reader has, or has as well. These are the arguments for choosing this crate over the incumbent, and none of them are the speed number the README leads with.

Capabilityfalcon_mdfThe fieldEvidence

Full capability catalogue

Every user-visible capability, grouped by what it is for. Filter to what you need. Shipped means implemented and covered by a test; gated means it needs a non-default cargo feature; partial means it works for a named subset; gap means absent, and the row says whether it refuses by name.

shipped implemented & tested gated behind a cargo feature partial named subset only gap absent
Show
CapabilityStatusWhat it doesEvidence

The viewer

A native egui desktop application, falcon, built on the same library. The viewer is a library plus a thin binary, so its logic is testable without an open window.

Panel / featureStatusWhat it showsEvidence

The browser viewer

falcon-mdf-wasm wraps the library with wasm-bindgen, and wasm/demo is a static site on top of it: no build step, no CDN, no plotting library, and all wasm in a worker so the page never blocks. Files stay on the visitor's machine. Deployed to GitHub Pages.

FeatureStatusWhat it doesEvidence

Gaps and refusals

Named so you can tell before you depend on it. A refusal reports itself through Mf4Error::Unsupported and the rest of the file still opens; a gap is simply not there. Nothing in the current tree fails Mf4File::open outright any more.

What is missingKindConsequenceAlso missing in

The defect pattern

Five defects of one shape have been found and fixed. Recording the shape is worth more than recording the fixes.

A size read from the file, used in arithmetic over bytes that are a different size

un_transpose divided the buffer by the column size rounding up, then discarded any byte whose destination landed past the end. Whenever the buffer length was not an exact multiple of the column size the output was reordered wrongly and carried a zero where a real byte belonged — no error, just a measurement with a hole in it. It affected zip type 1, transposed deflate, which is common, and it shipped.

Buffer / columnCorrectfalcon, before the fixResult
9 / 3t0 t3 t6 t1 t4 t7 t2 t5 t8sameagree
8 / 3t0 t2 t4 t1 t3 t5 t6 t7t0 t3 t6 t1 t4 t7 t2 t5wrong order
7 / 3t0 t2 t4 t1 t3 t5 t6t0 t3 t6 t1 t4 ZERO t2t5 lost
10 / 4t0 t2 t4 t6 t1 t3 t5 t7 t8 t9t0 t3 t6 t9 t1 t4 t7 ZERO t2 t5t8 lost

Why no test caught it. The round-trip tests transposed with falcon's own helper and un-transposed with its inverse, so they agreed with themselves rather than with the format. On exact multiples the two mappings coincide, and every test used an exact multiple. The replacement test states the expected bytes outright, so it is able to disagree — and against the old implementation it does.

The two siblings: record offsets advanced by a declared data-only size while the block emitted records interleaved with invalidation bytes, so later offsets walked into the previous block's tail and returned those bytes as samples; and a block-index path that concatenated invalidation bytes instead of interleaving them, parsing cleanly and decoding wrong. Every one of the three produced plausible wrong numbers rather than an error. Any new code that reads a length from a block and does arithmetic with it should be treated as guilty until a test says otherwise.

A fourth of the same shape surfaced while aligned multi-channel streaming was being built, and is worth recording because of how it was caught. The streaming path accumulated companion VLSD payloads per window rather than cumulatively, so on an unsorted J1939 log it returned different bytes than the eager read — plausible frame payloads, no error. The test that caught it asserts streamed output equals signals() over the entire corpus, and it failed only once the sample files were present: the agent that wrote the code reported it passing from a checkout where test_data/ is gitignored and the test had silently skipped. A corpus test that skips is not a test that passed, and a green run proves nothing about a checkout with no corpus in it.

A fifth arrived with the September hardening pass, in the most basic form the shape takes: a link table that declared a huge count but held only a few bytes of links. The parser allocated for the declared count and panicked on capacity; separately, an offset addition could overflow and saturating header arithmetic accepted an impossible layout. It now checks the bytes actually present before allocating and uses checked arithmetic. tests/link_bounds.rs failed four cases before the fix.

Against asammdf 8.7.2

The Python incumbent, and the benchmark this project has been measured against. Every row was checked against the installed package's source, not its documentation.

Show
CapabilityVerdictfalcon_mdfasammdfEvidence

Correctness, adjudicated against the standard rather than against a library

Full-length comparison over 75 files, 2,012 comparable channels, 9,233,007 values: 1,998 bit-for-bit identical, 1 within 1e-9, 13 genuine disagreements, and zero cases where falcon errored and asammdf succeeded. Of the 13, 9 are asammdf being wrong — VLSD record lengths, UTF-16 payload truncation, CANopen date/time fields, and a rational conversion whose result changes with the number of samples requested. Two are representation choices; two were unresolved and one of those, the value-range upper bound, has since been settled against falcon and fixed.

Against the Rust field

asammdf is the benchmark, but it is not the competition for a Rust crate. Versions, dates and download counts were re-fetched from the crates.io API.

CrateVersionLast publishDownloadsStanding against falcon_mdf

This reframes the goal

"Beat asammdf" measured this project against a Python library. Against the Rust field the picture is different: mf4-rs and mdf4-rs both write MDF4, and mf4-rs ships Python and wasm bindings plus HTTP byte-range sources. falcon has now matched the first two — it is on crates.io, has a pyo3 binding and a wasm-bindgen binding — so what mf4-rs still holds is a PyPI package and byte-range sources. What falcon leads the Rust field on: bounded-window streamed reading, a unified DBC + ARXML + LDF decoder with J1939 matching, typed non-f64 samples, the block-graph inspector, MDF 3.x and 4.x in one crate, all four logged buses, the export set, and a desktop and a browser viewer. None of the others advertise any of those.

Against everything else

The wider field, ordered by how often a team actually reaches for it. Rows marked unmeasured are characterisations from the tools' own documentation, not results produced here — they are here for orientation and should not be quoted as findings.

ToolKindWhat it is good atStanding against falcon_mdf

Performance record

The table most likely to be quoted back at us, so it carries the unflattering rows too. From benchmarks/COMPARISON.md, run 2026-09-27: 87 files, release build with LTO, a warm-up then the median of three runs, against asammdf 8.7.2 on CPython 3.14. Ratios compare equal work only — the three files where the libraries decode different sample counts, for reasons in the open questions, are excluded. Speed is size-dependent: never quote one number.

Scenevs asammdfReproducible hereNote

What is actually being raced

asammdf is not "Python" on the paths that matter. Its hot loops are a compiled C extension — blocks/cutils.c: sort_data_block, extract, get_vlsd_max_sample_size, invalidation-bit decode — and it decompresses through isal.isal_zlib (libisal) where available, with NumPy vectorisation above that. The rows where falcon lost were precisely the isal-accelerated compressed paths.

What changed. The September perf round switched flate2 to its zlib-rs backend, read compressed slices without an extra buffer, tiled un-transposition for cache locality, and loaded standard numeric widths directly. In a paired before/after run that took the 121.9 MiB transposed-deflate file from 1,733 ms to 1,078 ms (1.61×), and the 479.7 MiB uncompressed file 1.14×. The four 5 MiB J1939 logs moved 3–5% either way. The 1.76× below is on a different file from the old 0.81× row, so it shows the gap closed; it does not measure by how much. The 2026-09-27 rerun reads the same: 1.79× and 1.78×.

What is still thin. Both large fixtures are asammdf-written repetitions of J1939 logs, so they add size but not structural variety. The vendor-DZ files behind the 0.85–1.01× row are still not in the repository and were not re-run after zlib-rs.

Robustness record

1,049 mutated files per tool — truncation, corrupted lengths and links, bad block IDs, absurd counts, cyclic links. This is the strongest comparative result the project has.

ToolClean errorHangPanic / abortReturned data
falcon_mdf73703309
asammdf419665559
mdfreader627610361

Three hard failures against asammdf's 71 and mdfreader's 61, and falcon never hangs — cyclic block links, which hang both Python libraries, are rejected every time. A fresh sweep of 1,200 mutated files produces no panic, no abort and no hang; that sweep is only worth the paper it is written on because a deliberately crashing build was fed through the same harness first, to prove it reports a crash when one happens.

The claim is absolute, and absolute is still false

Both findings are now closed. The release-build panic — chunk size must be non-zero, from a 1,602-byte file with one 4-byte field zeroed — was a zero cn_bit_count reaching chunks_exact_mut. A zero-width channel now returns a named error, and a test asserts it fails rather than panics.

The subtler finding mattered more: cargo fuzz builds with debug assertions, so shift overflows in read_bits aborted and registered as crashes, while the release profile a consumer ships returned a wrong number silently. The unaligned path now accumulates in u128 and narrows only after shifting the field to bit zero, so the overflowing shift cannot occur — and two fuzz targets cover it: read_bits against an independent bit-by-bit oracle, runnable with --release, and a differential target comparing a debug and a release build on the same input.

Memory record

Peak memory scales with the largest data group, not the file. Under the memory-mapped backend the data is resident twice — as mapped pages and as the assembly buffer.

MeasurementResultStanding

The "roughly half the memory" claim for the buffered backend does not hold at any size that can be measured today. Choose it because another process may modify the file — not for a memory ratio. Against asammdf the large fixtures show falcon's whole-process peak 1.3–1.6× lower; those are process peaks under different read calls (falcon's whole read vs asammdf's get()), not isolated decoder allocations.

How this is verified

Research agents produced the raw inventories; every one was checked mechanically before it reached this page. Nothing here rests on an agent's word alone.

SourceCheck performedResult

What is next

The capability backlog is spent and the crate is published. What is left is release hygiene, very large files in the browser, and evidence from real files the crate did not write.

Open questions

QuestionStatus