Blog posts

2025

What We Have Built At Deep Render (So Far)

10 minute read

Published:

Deep Render was built from the ground up to develop a production ready codec. We focused on developing an AI codec with low computational complexity and high compression efficiency. We worked with a relentless product focus, efficient tooling and process driven research. After 4 years in the trenches, we have developed a codec that can achieve real-time encode and decode with over 45% BD rate saving w.r.t. SVT-AV1. In other words, the world’s first AI codec.

Future of AI based compression

7 minute read

Published:

Deep Render recently introduced the first AI codec into FFmpeg and VLC. This marks a significant step for AI codecs, making them readily available in tools widely used in the compression industry. Deep Render will continue to push the frontier of AI codecs by further improving compression performance, providing support for more hardware platforms and improving feature diversity.

2024

Solving AI based compression

5 minute read

Published:

Deep Render recently shipped the world’s first AI codec to its customers. This result follows two years of careful research, gruelling engineering, and relentless product focus. I’ll take some time to share some thoughts on the research and engineering that went into solving AI-based compression.

2021

Transformer and ViT dataflow notation

6 minute read

Published:

When reviewing the transformer and ViT literature, to get an intuitive understanding of the various model layers and how the input tokens are manipulated, I found it helpful to map out the data flow through the model in matrix notation. I couldn’t find this anywhere so I thought I’d share it.

Diffusion Decoder based Compression

11 minute read

Published:

I’m always super interested in expanding the fronteir of learned compression. The current methods we use are derived from VAEs (equivalence shown here) and use GANs to improve preceptual quality but in this blog I explore how diffusion models could be used for compression.

Proof: MSE makes GANs locally stable

7 minute read

Published:

Generative Adversarial Networks are notoriously unstable due to issues such as mode collapse and training divergence. However, in AI compression, the adversarial training is generally stable and reliable without the neccessary tricks such as gradient penalty and adding noise to our input samples. I found this intriguing so I set out to explore why.

2019

Cross-Entropy, KL and NLL are the same objective in AI compression

2 minute read

Published:

In end-to-end learned compression, we need to model the distribution of data, so we can entropy encode it. The training loss we use is the cross-entropy between the unknown true data distribution and our model distribution which in a simplest case is a fully factorized distribution of standard normals. There is often some confusion about the objective used in compression so I thought I’d use this post to clarify it. I’ll show that for data sampled from the true distribution \(p\) and a parametric model \(q_\theta\) we learn, the cross-entropy, the Kullback–Leibler divergence and the negative log-likelihood are optimization-equivalent objectives. In practice this means we are doing maximum likelihood estimation (MLE) and simultaneously minimizing expected code length.