Blogs · Aras' website

Compact Normal Storage for small g-buffers

Posted on Aug 4, 2009

I’ve been experimenting with compact storage of view space normals for small g-buffers. Think about storing depth and normal in a single 8 bit/channel RGBA texture.

Here are my findings - with error visualization and shader performance numbers for some GPUs.

If you know any other method to encode/store normals in a compact way, please let me know!

Encoding floats to RGBA - the final?

Posted on Jul 30, 2009

The saga continues! In short, I need to pack a floating point number in [0..1) range into several channels of 8 bit/channel render texture. My previous approach is not ideal.

Turns out some folks have figured out an approach that finally seems to work.

Here it is for my own reference:

gamedev.net forum post by gjaegy
Suggestion right there on my previous blog post comments
Repost gamerendering blog
Repost on gamedev.net forums again.

So here’s the proper way:

inline float4 EncodeFloatRGBA( float v ) {
  float4 enc = float4(1.0, 255.0, 65025.0, 16581375.0) * v;
  enc = frac(enc);
  enc -= enc.yzww * float4(1.0/255.0,1.0/255.0,1.0/255.0,0.0);
  return enc;
}
inline float DecodeFloatRGBA( float4 rgba ) {
  return dot( rgba, float4(1.0, 1/255.0, 1/65025.0, 1/16581375.0) );
}

That is, the difference from the previous approach is that the “magic” (read: hardware dependent) bias is replaced with subtracting next component’s encoded value from the previous component’s encoded value.

Implementing fixed function T&L in vertex shaders

Posted on Jun 9, 2009

Almost half a year ago I was wondering how to implement T&L; in vertex shaders.

Well, finally I implemented it for upcoming Unity 2.6. I wrote some sort of a technical report here.

In short, I’m combining assembly fragments and doing simple temporary register allocation, which seems to work quite well. Performance is very similar to using fixed function (I know it’s implemented as vertex shaders internally by the runtime/driver) on several different cards I tried (Radeon HD 3xxx, GeForce 8xxx, Intel GMA 950).

What was unexpected: the most complex piece is not the vertex lighting! Most complexity is in how to route/generate texture coordinates and transform them. Huge combination explosion there.

Otherwise - I like! Here’s a link to the article again.

Shaders must die, part 3

Posted on May 10, 2009

Continuing the series (see Part 1, Part 2)…

Got different lighting models (BRDFs) working. Without further ado, code snippets that produce real actual working shaders that work with lights & shadows and whatnot:

Simple Lambert (single color):

 Properties
     Color _Color
 EndProperties
 Surface
     o.Albedo = _Color;
 EndSurface
 Lighting Lambert

Let’s add a texture:

 Properties
     2D _MainTex
     Color _Color
 EndProperties
 Surface
     o.Albedo = SAMPLE(_MainTex) * _Color;
 EndSurface
 Lighting Lambert

Change light model to Half-Lambert (a.k.a. wrapped diffuse):

 // ...everything the same
 Lighting HalfLambert

Blinn-Phong, with constant exponent & constant specular color, modulated by gloss map in main texture’s alpha:

 Properties
     2D _MainTex
     Color _Color
     Color _SpecColor
     Float _Exponent
 EndProperties
 Surface
     half4 col = SAMPLE(_MainTex);
     o.Albedo = col * _Color;
     o.Specular = _SpecColor.rgb * col.a;
     o.Exponent = _Exponent;
 EndSurface
 Lighting BlinnPhong

The same Blinn-Phong, with added normal map:

 Properties
     2D _MainTex
     2D _BumpMap
     Color _Color
     Color _SpecColor
     Float _Exponent
 EndProperties
 Surface
     half4 col = SAMPLE(_MainTex);
     o.Albedo = col * _Color;
     o.Specular = _SpecColor.rgb * col.a;
     o.Exponent = _Exponent;
     o.Normal = SAMPLE_NORMAL(_BumpMap);
 EndSurface
 Lighting BlinnPhong

I also made an illustrative-style BRDF (see Illustrative Rendering in Team Fortress 2), but that only requires above sample to have “Lighting TF2” at the end.

Another thing I tried is surface that has Albedo dependent on a viewing angle, similar to Layered Car Paint Shader. It works:

 Properties
     2D _MainTex
     2D _BumpMap
     2D _SparkleTex
     Float _Sparkle
     Color _PrimaryColor
     Color _HighlightColor
 EndProperties
 Surface
     half4 main = SAMPLE(_MainTex);
     half3 normal  = SAMPLE_NORMAL(_BumpMap);
     half3 normalN = normalize(SAMPLE_NORMAL(_SparkleTex));
     half3 ns = normalize (normal + normalN * _Sparkle);
     half3 nss = normalize (normal + normalN);
     i.viewDir = normalize(i.viewDir);
     half nsv = max(0,dot(ns, i.viewDir));
     half3 c0 = _PrimaryColor.rgb;
     half3 c2 = _HighlightColor.rgb;
     half3 c1 = c2 * 0.5;
     half3 cs = c2 * 0.4;    
     half3 tone =
         c0 * nsv +
         c1 * (nsv*nsv) +
         c2 * (nsv*nsv*nsv*nsv) +
         cs * pow(saturate(dot(nss,i.viewDir)), 32);
     main.rgb *= tone;
     o.Albedo = main;
     o.Normal = normal;
 EndSurface
 Lighting Lambert

Up next:

How and where emissive terms should be placed. I cautiously omitted all emissive terms from the above examples (so my layered car shader is without reflections right now).
Where should things like rim lighting go? I’m not sure if it’s a surface property (increasing albedo/emission with angle) or a lighting property (a back light).

My impressions so far:

I like that I don’t have to write down vertex-to-fragment structures or the vertex shader. In most cases all the vertex shader does is transform stuff and pass it down to later stages, plus occasional computations that are linear over the triangle. No good reason to write it by hand.
I like that the above shaders do not deal with how the rendering is actually done. For Unity’s case, I’m compiling them into single pass per light forward renderer, but they should just work with multiple lights per pass, deferred etc. Of course, that still has to be proven!

So far so good.

Series index: Shaders must die, Part 1, Part 2, Part 3.

Shaders must die, part 2

Posted on May 7, 2009

I started playing around with the idea of “shaders must die”. I’m experimenting with extracting “surface shaders” for now.

Right now my experimental pipeline is:

Write a surface shader file
Perl script transforms it into Unity 2.x shader file
Which in turn is compiled by Unity into all lighting/shadows permutations, for D3D9 and OpenGL backends. Cg is used for actual shader compilation.

I have very simple cases working. For example:

 Properties
     2D _MainTex
 EndProperties
 Surface
     o.Albedo = SAMPLE(_MainTex);
 EndSurface

This is a “no bullshit” source code for a simple Diffuse (Lambertian) shader, 87 bytes of text.

The Perl script produces a Unity 2.x shader. This will be long, but bear with me - I’m trying to show how much stuff has to be written right now, when we’re operating on vertex/pixel shader level. See Attenuation and Shadows for Pixel Lights in Unity docs for how this system works.

 Shader "ShaderNinja/Diffuse" {
 Properties {
   _MainTex ("_MainTex", 2D) = "" {}
 }
 SubShader {
   Tags { "RenderType"="Opaque" }
   LOD 200
   Blend AppSrcAdd AppDstAdd
   Fog { Color [_AddFog] }
   Pass {
     Tags { "LightMode"="PixelOrNone" }
 CGPROGRAM
 #pragma fragment frag
 #pragma fragmentoption ARB_fog_exp2
 #pragma fragmentoption ARB_precision_hint_fastest
 #include "UnityCG.cginc"
 uniform sampler2D _MainTex;
 struct v2f {
     float2 uv_MainTex : TEXCOORD0;
 };
 struct f2l {
     half4 Albedo;
 };
 half4 frag (v2f i) : COLOR0 {
     f2l o;
     o.Albedo = tex2D(_MainTex,i.uv_MainTex);
     return o.Albedo * _PPLAmbient * 2.0;
 }
 ENDCG
   }
   Pass {
     Tags { "LightMode"="Pixel" }
 CGPROGRAM
 #pragma vertex vert
 #pragma fragment frag
 #pragma multi_compile_builtin
 #pragma fragmentoption ARB_fog_exp2
 #pragma fragmentoption ARB_precision_hint_fastest
 #include "UnityCG.cginc"
 #include "AutoLight.cginc"
 struct v2f {
     V2F_POS_FOG;
     LIGHTING_COORDS
     float2 uv_MainTex;
     float3 normal;
     float3 lightDir;
 };
 uniform float4 _MainTex_ST;
 v2f vert (appdata_tan v) {
     v2f o;
     PositionFog( v.vertex, o.pos, o.fog );
     o.uv_MainTex = TRANSFORM_TEX(v.texcoord, _MainTex);
     o.normal = v.normal;
     o.lightDir = ObjSpaceLightDir(v.vertex);
     TRANSFER_VERTEX_TO_FRAGMENT(o);
     return o;
 }
 uniform sampler2D _MainTex;
 struct f2l {
     half4 Albedo;
     half3 Normal;
 };
 half4 frag (v2f i) : COLOR0 {
     f2l o;
     o.Normal = i.normal;
     o.Albedo = tex2D(_MainTex,i.uv_MainTex);
     return DiffuseLight (i.lightDir, o.Normal, o.Albedo, LIGHT_ATTENUATION(i));
 }
 ENDCG
   }
 }
 Fallback "VertexLit"
 }

Phew, that is quite some typing to get simple diffuse shader (1607 bytes)! Well, at least all the lighting/shadow combinations are handled by Unity macros here. When Unity takes this shader and compiles into all permutations, it results in 58 kilobytes of shader assembly (D3D9 + OpenGL, 17 light/shadow combinations).

Let’s try something slightly different: bumpmapped, with a detail texture:

 Properties
     2D _MainTex
     2D _Detail
     2D _BumpMap
 EndProperties
 Surface
     o.Albedo = SAMPLE(_MainTex) * SAMPLE(_Detail) * 2.0;
     o.Normal = SAMPLE_NORMAL(_BumpMap);
 EndSurface

This is 173 bytes of text. Generated Unity shader is 2098 bytes, which compiles into 74 kilobytes of shader assembly.

In this case, the processing script detects that surface shader modifies normal per pixel, and does the necessary tangent space light transformations. It all just works!

So this is where I am now. Next up: detect which lighting model to use based on surface parameters (right now it always uses Lambertian). Fun!