CSV, MATLAB, TDMS, Parquet or HDF5?
Five common formats for measurement data, what each is good at, and which to pick for your workflow.
Once a test is done, the data has to go somewhere your analysis tools can read. The format you pick decides how big the files are, how fast they open and whether the context (units, sample rate, device, notes) survives the trip. Here is how the five formats engineers most often ask for compare.
| Format | Best for | Size and speed | Metadata | Opens in |
|---|---|---|---|---|
| CSV | Sharing small datasets with anyone | Large and slow to parse | Header row only, by convention | Everything |
| MATLAB .mat | Teams that analyze in MATLAB | Compact and fast | Variables and structs | MATLAB, Python (scipy for v5, h5py for v7.3) |
| NI TDMS | Labs using LabVIEW or DIAdem | Compact and fast | Groups, channels and properties | NI tools, Python (npTDMS), Excel plug-in |
| Parquet | Data pipelines and dataframes | Columnar and compressed | Typed schema, key-value metadata | pandas, Polars, DuckDB, Spark |
| HDF5 | Large, long or multi-device recordings | Chunked, compressible, fast slicing | Rich, hierarchical attributes | Python (h5py), MATLAB, Julia, many others |
CSV
CSV is the universal fallback: every tool opens it and you can read it in a text editor. It is also the least efficient. Numbers stored as text take several times the space of binary, parsing is slow, and there is no standard place for units, sample rate or device information. Excel stops at 1,048,576 rows, so a few minutes of multi-channel data at a kilohertz already overflows it. Use CSV for small exports and for handing data to people who just need to open it.
MATLAB .mat
If your analysis lives in MATLAB, a .mat file loads as ready-to-use variables. Note the versions: the older v5 format opens in Python with scipy.io.loadmat, while v7.3 files are HDF5 underneath and need h5py or a similar library.
NI TDMS
TDMS is National Instruments’ format and the native language of LabVIEW and DIAdem. It stores data as groups and channels with properties attached, which maps neatly onto DAQ recordings. Outside NI’s tools, the npTDMS Python package reads it well. Pick TDMS when the people analyzing your data already work in NI software.
Apache Parquet
Parquet is a columnar format from the data engineering world. It compresses well, loads selected columns quickly and is the natural input for pandas, Polars, DuckDB and Spark. It suits test data that will be queried across many runs, for example every unit tested on a production line this month.
HDF5
HDF5 is built for large numeric arrays. Data is chunked and optionally compressed, so you can read a ten-second slice of an hour-long recording without loading the whole file, and the hierarchy of groups and attributes keeps context next to the data. It is the strongest general-purpose choice for long or multi-device recordings, and nearly every scientific language can read it.
So which one?
- Sending a few thousand rows to a colleague: CSV.
- Analysis happens in MATLAB: .mat.
- Analysis happens in LabVIEW or DIAdem: TDMS.
- Results feed a database or dataframe pipeline: Parquet.
- Long recordings, many channels, or you are not sure yet: HDF5.
How DAQAtlas stores and exports data
DAQAtlas, which is in development, records to its own .edaq format while a test runs: a small JSON header with the device, channels, units, rate and start time, followed by raw samples. Because the frame count comes from the file size, a recording stays readable even if the app quits mid-run. From review you can export everything, the visible range or the span between two cursors to CSV, MATLAB, TDMS, Parquet or HDF5, with optional decimation and wall-clock timestamps. Each export format is checked in our test suite by reading it back with scipy, npTDMS, pyarrow and h5py.
Join the beta to try it with your own data.