Help Center

FAQ file formats future proofing

Are there any limitations to what files you can normalize?

Due to processing limitations, we only process files smaller than 100 MB. If this poses a challenge, we may be able to offer alternative solutions.

How do you define what formats to use?

At Piql, we categorize file formats into two types: Preferred and Accepted formats. Preferred formats are simple formats that are easier to interpret and work with.

For example, a normalized TIFF file stores each pixel as a single point. While this approach takes up more space than using compression techniques, it ensures better long-term preservation.

Accepted formats are widely used formats expected to remain accessible for many years, such as JPEG and Microsoft Office Open formats.

PDFs receive special treatment. Since PDF is a complex format, we use VeraPDF to validate if it's archivable. We classify PDF/A formats as Preferred, while other PDF formats are considered Accepted.

What about other formats?

When future-proofing is enabled in your package, we attempt to convert your files to a Preferred format. For instance, if you upload a PNG image file, we'll convert it to a normalized TIFF file. For file formats we don't support, the file will be labeled "Unsupported." You can contact us to discuss possible updates to support your specific format.

How are the conversions actually done?

We employ several specialized tools for conversion: FFMPEG for audio files, ImageMagick for images, and LibreOffice for converting office documents to PDF. These are best-effort conversions that should be verified, as a general solution cannot match the quality of format-specific customized conversion.

Converted files are typically larger than originals. For image conversion, we remove compression and save each pixel individually. By default, we preserve your original file, though there's an option to remove it after successful conversion.

How do you detect formats?

We use the 'file' tool in Linux to determine each file's MIME type. This means we can correctly identify file types even when they're mislabeled—for example, detecting a JPEG file that was incorrectly named as a TIFF.

However, this detection isn't foolproof. Since there's no universal standard for file type detection, some formats—particularly those based on or compatible with other formats—can be difficult or impossible to distinguish from one another.

On this page