Showing posts with label Y' Cb Cr. Show all posts
Showing posts with label Y' Cb Cr. Show all posts

Tuesday, November 5, 2013

Color Subsampling Notation

It's time to return to this blog and I'm getting started by taking some of the most popular posts and updating them and featuring content that still sparks conversations like one I had today with a friend and colleague that inspired me to revisit this post (I originally posted this topic in 2008).

Color Subsampling gets confused with color precision (8 bits per channel, 10 bits per channel, etc.) and color channels (Does 4:2:0 mean there isn't ANY Cr channel samples?).

In actuality, the number is really best characterized as a ratio. (click on the chart for a large visual)


“4” in the first slot is easiest to think of as representing the baseline of four pixels (and these ratios only apply to digital video signals).  The first number represents the first channel of the RGB or Y'CbCr group.

The second number and third number are frequently thought to represent the remaining two channels, but actually the second number refers to the sampling frequency of both the second and third channels horizontally and the third number was originally intended to indicate the sampling frequency of both vertically, though the system was developed without really considering vertical subsampling systems like 4:2:0. 

In the current system, the third number is either the same as the second number as in 4:2:2 and 4:1:1 indicating no vertical subsampling…all the vertical color difference samples are there in each column that has a horizontal color difference sample. In ratios where the third number is zero, the “0” indicates that there is a 2:1 vertical subsample in addition to the horizontal color difference subsample.



4:4:4
A designation of 4:4:4 would mean that there is a discreet sample for each of three color channels making up the signal for every pixel. While this could apply to either RGB or Y'CbCr used for video, 4:4:4 would most often be seen with an RGB signal, but 4:4:4 could refer to a Y'CbCr color sampling scheme 

RGB does not subsample one color channel in relation to another, so 4:2:2 (or 4:1:1, etc...) would never refer to RGB.

4:2:2
This number is most prevalent in high-end video formats and refers to a discrete sample for Y’ (luma) on every pixel and samples for each color difference signal is sampled at one value for every two pixels. While in theory this sounds like the elimination of a lot of information (a third actually) compared to 4:4:4, the human eye prioritizes the detail in the luma portion of the image and most humans would be hard pressed to see the difference between a color Y’ CB CR image in 4:4:4 and one in 4:2:2. In fact, 4:2:2 is good enough that most video types that are designated as “uncompressed” are actually color subsampled at 4:2:2.

4:1:1
NTSC DV introduced us to this aggressive, lossy color subsampling scheme. For every four Y’ samples horizontally, there is only one sample for Cb and Cr.  This creates a 4x1 four pixel horizontal “block” with common color difference values, though each pixel has a discreet Y’ value so the pixels aren’t identical. 

While DV footage was used extensively, even in broadcasting, it can be a challenge for special effects and compositing as chroma keying and green and blue screen work requires a lot of subtle tonal variations to recognize irregular vertical edges. Canopus and Matrox each created custom methods of decode for DV to attempt to improve (effectively interpolating to 4:2:2) the four pixel horizontal spread for better keying, and many software keyers have similar measures in place. 

It's interesting to note that even though 4:2:0 subsampling is thought by many to be somewhat inferior to 4:1:1, 4:2:0 (compression set aside from color subsample for a moment) can actually be slightly easier to composite or key as there is only one pixel of interpolated value in either the vertical or horizontal direction, while 4:1:1 interpolates 3 pixel values horizontally.

4:2:0
PAL DV users and anyone who outputs to MPEG has seen this number. Many people find it confusing at first as it appears the notation as a Y’ sample for each pixel, a Cb sample for every two pixels, and no samples whatsoever for Cr.  In reality, there are the same number of color difference samples as NTSC DV with the pixels arranged differently. 

Also confusing: all the color difference sample sites for the various approaches to 4:2:0 are not standard. (see chart) JPEG/MPEG-1 structures the samples so that they’re sited in the center of the four pixel block. MPEG-2 sites the samples between pixels vertically, and PAL DV sites the difference samples on alternating lines. Even with the color difference samples sited differently for different applications of 4:2:0, you could say there are still four pixel blocks that net out to the same amount of color difference samples as 4:1:1 and simply picture these 4 pixel “blocks” as square (2x2) instead of a horizontal line (4x1) like NTSC DV’s 4:1:1.

4:2:2:4, 4:4:4:4
As if all this isn’t complicated enough…you could add a number. 4:2:2:4 or 4:4:4:4 refer to 4:2:2 or 4:4:4 color sampling with the addition of an alpha channel for keying purposes. The fourth channel would carry an 8 bit or 10 bit (depending on the image format) grayscale map indicating relative transparency of each pixel in the image. The alpha number is always the same as the Y' sample.

3:1:1
This ratio appeared during the period of HDCAM's introduction.   Playback is 1920x1080, but actually record 1440x1080 to tape. In my opinion the most confusing aspect is not so much that there is a different baseline number, but whether or not that number is a proportion of “4” in itself as 1440/1920 is 3 of 4. 

The interpolation to 1920x1080 4:2:2 (this is how the manufacturer presents the specs on the playout picture) and the color difference subsampling ratio of 3:1:1 are separate issues and their mathematical scale to full raster 1920x1080 is most likely coincidental. 3 equates to 1440 horizontal Y’ samples and 1 is a ratio to 3 designating 480 horizontal color difference samples. This notation is NOT on the chart as it does not exist anywhere but in file storage, and the end user can only access HDCAM footage as 4:2:2 SDI output without a proprietary post solution.

As we continue to see new formats and frame sizes, we'll continue to see new approaches to storing and encoding images, but the color subsampling notations will likely stay in place for the foreseeable future.

TimK

Monday, December 8, 2008

Color Subsample Notation


The numbers get tossed around with impunity these days. The first number is usually 4 and the closer the next two are to 4 the better…right? Well, while that statement is basically true, there’s a lot more to it than just that. The number is really a ratio. (click on the chart for a large visual)

“4” in the first slot is meant to represent the baseline of four pixels and these ratios only apply to digital video signals. The physical arrangement of the four pixels in question isn’t really referred to in a standard way any longer but originally it was supposed to represent 4 horizontal pixels. The second number and third number are frequently, but erroneously assumed to represent the relative sampling value for each of the color difference channels. Actually the second number refers to the sampling frequency of both difference signals horizontally and the third number was originally intended to indicate the sampling frequency of both difference signals vertically, though the system was developed without really considering vertical subsampling systems like 4:2:0. In the current system, the third number is either the same as the second number as in 4:2:2 and 4:1:1 indicating no vertical subsampling…all the vertical color difference samples are there in each column that has a horizontal color difference sample. In ratios where the third number is zero, the “0” indicates that there is a 2:1 vertical subsample in addition to the horizontal color difference subsample.
4:4:4
A designation of 4:4:4 would mean that there is a discreet sample for each of three color channels making up the signal for each pixel. While this could apply to either RGB or one of the color difference color spaces used for video, 4:4:4 would most often be seen with an RGB signal. Even though 4:4:4 could refer to a Y'CbCr color sample, RGB does not subsample one color channel in relation to another, so 4:2:2 (or 4:1:1, etc...) would never refer to RGB.
4:2:2
This number is most prevalent in high-end video formats and refers to a discrete sample for Y’ on every pixel and samples for each color difference signal is sampled at one value for every two pixels. While in theory this sounds like the elimination of a lot of information (a third actually) compared to 4:4:4, the human eye prioritizes the detail in the luma portion of the image and most humans would be hard pressed to see the difference between a color Y’ CB CR image in 4:4:4 and one in 4:2:2. In fact, 4:2:2 is good enough that most video types that are designated as “uncompressed” are actually color sampled at 4:2:2.
4:1:1
Most users of NTSC DV are familiar with this color sampling scheme. For every four Y’ samples, there is only one sample for CB and CR. This creates a 4x1 four pixel horizontal “block” with common color difference values, though each pixel has a discreet Y’ value so the pixels aren’t identical. While DV footage is used extensively, even in broadcasting, it can be a challenge for special effects and compositing as chroma keying and green and blue screen work requires a lot of subtle tonal variations to create smooth irregular vertical edges. Canopus and Matrox each created custom methods of decode for DV to attempt to better interpolate the four pixel horizontal spread for better keying, and many software keyers have similar measures in place. It's intersting to note that even though 4:2:0 subsampling is thought by many to be somewhat inferior to 4:1:1, 4:2:0 (compression set aside from color subsample for a moment) can actually be slightly easier to composite as there is only one pixel of interpolated value in either the vertical or horizontal direction, while 4:1:1 interpolates 3 pixel values horizontally.
4:2:0
PAL DV users and anyone who outputs to MPEG has seen this number. Many who may initially interpret the notation as a Y’ sample for each pixel, a CB sample for every two pixels, and no samples whatsoever for CR can find it confusing. In reality, there are the same number of color difference samples as NTSC DV with the pixels arranged differently. Also confusing: all the color difference sample sites for the various approaches to 4:2:0 are not standard. (see chart) JPEG/MPEG-1 structures the samples so that they’re sited in the center of the four pixel block. MPEG-2 sites the samples between pixels vertically, and PAL DV sites the difference samples on alternating lines. Even with the color difference samples sited differently for different applications of 4:2:0, you could say there are still four pixel blocks that net out to the same amount of color difference samples as 4:1:1 and simply picture these 4 pixel “blocks” as square (2x2) instead of a horizontal line (4x1) like NTSC DV’s 4:1:1.
4:2:2:4, 4:4:4:4
As if all this isn’t complicated enough…you could add a number. 4:2:2:4 or 4:4:4:4 refer to 4:2:2 or 4:4:4 color sampling with the addition of an alpha channel for keying purposes. The fourth channel would carry an 8 bit or 10 bit (depending on the image format) grayscale map indicating relative transparency of each pixel in the image. The alpha number is always the same as the Y' sample.
3:1:1
As if the strange way the second and third value seem to fall in these ratios isn’t confusing enough…now we see a ratio where the first number has changed. This ratio appears when referring to HDCAM pictures which on playback are 1920x1080, but actually record 1440x1080 to tape. In my opinion the most confusing aspect is not so much that there is a different baseline number, but whether or not that number is a proportion of “4” in itself as 1440/1920 is 3 of 4. I suspect the interpolation to 1920x1080 4:2:2 (this is how the manufacturer presents the specs on the playout picture) and the color difference subsampling ratio of 3:1:1 are separate issues and their mathematical scale to full raster 1920x1080 is most likely coincidental. 3 equates to 1440 horizontal Y’ samples and 1 is a ratio to 3 designating 480 horizontal color difference samples. This notation is NOT on the chart as it does not exist anywhere but in file storage, and the end user can only access HDCAM footage as 4:2:2 SDI output without a proprietary post solution.

TimK

Saturday, December 6, 2008

As Promised... General Colorspace Information

We hear a lot about color space. RGB, Y'U'V', Y'/R-Y/B-Y, Y’/Cr/Cb…it ranges from logical to alphabet soup. In the end, the color in our images is made from three component signals combining to create a color image. What those three component signals are…well, that depends. (There are colorspace definitions outside of the colorspaces I've listed here, of course. I just tried to include the colorspaces most media production pros will encounter...)

RGB
The easiest place to start might be RGB. It will be review for almost anyone with any computer background as RGB logically stands for “Red/Green/Blue”. RGB color images are produced by combining three channels of image information, each containing the intensity information for each individual color, Red, Green, and Blue. This colorspace is typically used in computer displays and projectors. It’s used in the creation of the visual content developed for video games or the web. Conventional analog and digital video has traditionally not used RGB color space, causing many long-time professionals to discount the general accuracy of the video your NLE shows you on your computer desktop monitor. However, with the growing use of large LCD and Plasma displays RGB is becoming at least one of the display methods you may want to (or be forced to…) employ through the post process.



HSL
Hue, Saturation, Luminance” is found most often inside computer software for image manipulation or video color correction. It is fairly simple to picture as a color wheel where the color red is in one direction, moving around the wheel to yellow, green, cyan, blue, magenta, and back to red. While the color order would seem to directly mimic the vectorscope display, in HSL the individual colors are evenly spaced like numbers on a clock face while a vectorscope is showing a color difference signal where the positions of the primary and secondary legal colors are not equidistant.
However, the vectorscope model is indicative of how the hue and saturation portions of an HSL color wheel work. The “compass” direction from center indicates the hue and the distance from center indicates saturation. Also similar is the way the HSL color wheel doesn’t display luminance changes, but the brightness is adjusted via each color channel somewhat similar to an RGB model.

CMYK
For those of you who have dealt with this color space strictly as an irritant when you receive your client’s logo art from their printer and forget to convert it before spending 10 minutes trying to figure out why it won't load into After Effects, we’ll just touch briefly on what it is. Printing ink is based on reflected light whereas the RGB color model is constructed to create an image from a light source.
RGB is based on the color we perceive when we look directly at a light source (spotlight, television screen, LCD monitor…etc.) and is referred to as “Additive Color.” The primary colors are Red, Green and Blue while the secondary colors are Cyan, Magenta, and Yellow. When all colors are combined in an additive color environment (all colors of light are 'on'), you get white, when all color is gone (lights off), you get black.
CMYK turns everything on its head, as it's based on “Subtractive Color.” This is color we perceive once light has bounced off of something. Usually the color we see is the only color not absorbed by the object. So the car isn’t really green, the surface of the car absorbs most of the spectrum of light that hits it, but rejects green light and it bounces off, becoming visible to us. Since we are turning everything around, you can picture combining all the colors in a paint store (or as a kid, in your water color tray...or on the living room carpeting...) and getting something that is usually undesirably close to black. Removing all pigment from paint gives you…white. Just to cap off the confusion, the primary colors are Cyan, Magenta and Yellow, hence the “CMY” with “K” indicating black ink. Of course since all else is opposite, Red, Green and Blue are the secondary colors in subtractive color.

Color Difference Systems
I explained my rather unconventional way of visualizing Color Difference systems in general. (see the explanation and illustration from the post "Y's and Whats..." below). What follows is a description of the various common designations we all encounter in the field and their intended definition.

Y’UV
I’ve even spouted this designation off when I am referring to all color difference, non-RGB television color space, but it’s just laziness on my part as it really isn’t correct.
It does refer to color difference video color space, but it specifically refers to the color difference signals B’-Y’ and R’-Y’ after scaling used in an intermediate step in the creation of an analog composite NTSC or PAL signal. Though “Y’UV is frequently misused in referring to digital video, and particularly incorrect in cases where it is paired with a subsampling designation such as “4:2:2”, as it has nothing to do with component video or digital video.

Y’ PB PR
Y’ PB PR is the description for the coding for analog component video. PB and PR correspond to the B’-Y’ and R’-Y’ signal after scaling to this standard. In the case of analog, digital color sampling ratios like “4:2:2” don’t apply, but in this system the color difference components are lowpass filtered to approximately half the luma bandwidth.

Y’ CB CR
Digital component video changes the designation somewhat. Y’ CB CR references the scaling method used in digital video files on disk or tape. While the analog component color difference signals would be lowpass filtered to preserve bandwidth, their digital counterparts would instead be subsampled if efficiency was required, creating the basis for the typical ratios expressed with numbers like “4:2:2” or “4:2:0”. (More on that in later posts.)

Y’ I Q
An obsolete system established in the early days of color television. What was intriguing about this system is the normal vectorscope color orientation was tilted 33 degrees, and the “I” and “Q” axes (which correspond to the ‘U’ and ‘V’ components, but the axes are actually switched) can still be seen on some vectorscopes to this day in that configuration.
There's some stuff to keep your mind working over the weekend...
TimK