CSV, MATLAB, TDMS, Parquet or HDF5?

Five common formats for measurement data, what each is good at, and which to pick for your workflow.

Britt Espinosa··7 min read

Once a test is done, the data has to go somewhere your analysis tools can read. The format you pick decides how big the files are, how fast they open and whether the context (units, sample rate, device, notes) survives the trip. Here is how the five formats engineers most often ask for compare.

FormatBest forSize and speedMetadataOpens in
CSVSharing small datasets with anyoneLarge and slow to parseHeader row only, by conventionEverything
MATLAB .matTeams that analyze in MATLABCompact and fastVariables and structsMATLAB, Python (scipy for v5, h5py for v7.3)
NI TDMSLabs using LabVIEW or DIAdemCompact and fastGroups, channels and propertiesNI tools, Python (npTDMS), Excel plug-in
ParquetData pipelines and dataframesColumnar and compressedTyped schema, key-value metadatapandas, Polars, DuckDB, Spark
HDF5Large, long or multi-device recordingsChunked, compressible, fast slicingRich, hierarchical attributesPython (h5py), MATLAB, Julia, many others

CSV

CSV is the universal fallback: every tool opens it and you can read it in a text editor. It is also the least efficient. Numbers stored as text take several times the space of binary, parsing is slow, and there is no standard place for units, sample rate or device information. Excel stops at 1,048,576 rows, so a few minutes of multi-channel data at a kilohertz already overflows it. Use CSV for small exports and for handing data to people who just need to open it.

MATLAB .mat

If your analysis lives in MATLAB, a .mat file loads as ready-to-use variables. Note the versions: the older v5 format opens in Python with scipy.io.loadmat, while v7.3 files are HDF5 underneath and need h5py or a similar library.

NI TDMS

TDMS is National Instruments’ format and the native language of LabVIEW and DIAdem. It stores data as groups and channels with properties attached, which maps neatly onto DAQ recordings. Outside NI’s tools, the npTDMS Python package reads it well. Pick TDMS when the people analyzing your data already work in NI software.

Apache Parquet

Parquet is a columnar format from the data engineering world. It compresses well, loads selected columns quickly and is the natural input for pandas, Polars, DuckDB and Spark. It suits test data that will be queried across many runs, for example every unit tested on a production line this month.

HDF5

HDF5 is built for large numeric arrays. Data is chunked and optionally compressed, so you can read a ten-second slice of an hour-long recording without loading the whole file, and the hierarchy of groups and attributes keeps context next to the data. It is the strongest general-purpose choice for long or multi-device recordings, and nearly every scientific language can read it.

So which one?

  • Sending a few thousand rows to a colleague: CSV.
  • Analysis happens in MATLAB: .mat.
  • Analysis happens in LabVIEW or DIAdem: TDMS.
  • Results feed a database or dataframe pipeline: Parquet.
  • Long recordings, many channels, or you are not sure yet: HDF5.

How DAQAtlas stores and exports data

DAQAtlas, which is in development, records to its own .edaq format while a test runs: a small JSON header with the device, channels, units, rate and start time, followed by raw samples. Because the frame count comes from the file size, a recording stays readable even if the app quits mid-run. From review you can export everything, the visible range or the span between two cursors to CSV, MATLAB, TDMS, Parquet or HDF5, with optional decimation and wall-clock timestamps. Each export format is checked in our test suite by reading it back with scipy, npTDMS, pyarrow and h5py.

Join the beta to try it with your own data.