Friday, August 14, 2015

Building GLEW and GLFW libraries for static linking in Visual Studio

It seems the prebuilt binaries for GLEW and GLFW are not really suitable for static linking when using static (i.e. the non-DLL versions) of the VC++ runtime libraries.

Here's a short how-to for compiling GLEW and GLFW for this usage scenario:

1.) Building GLEW

  • download and extract GLEW source code package from http://glew.sourceforge.net/
  • goto /build/vc10 (or vc12)
  • open glew.sln
  • select desired build configuration (debug or release)
  • select glew_static project
  • set C/C++ Code Generation options for glew_static project: Select correct runtime library
  • Build project
  • Result: glew32s(d).lib
2.) Building GLFW
  • download and extract GLFW source code package from http://www.glfw.org/download.html
  • download and extract CMake win .zip from http://www.cmake.org/download/
  • run CMake GUI (i.e. bin/cmake-gui.exe)
  • select GLFW source root directory
  • choose output directory ("Where to build the binaries"), can be same as above
  • click "Configure"
  • select compiler
  • deselect "USE_MSVC_RUNTIME_LIBRARY_DLL" option
  • make sure "BUILD_SHARED_LIBS" is also disabled
  • click "Configure" again
  • click "Generate"
  • open generated GLFW.sln (located in output directory as specified above)
  • select desired build configuration (debug or release)
  • set C/C++ Code Generation options for glfw project: Select correct runtime library
  • Build project
  • Result: glfw3.lib


Monday, March 31, 2014

NVScene demo released

I participated in the PC demo competition at NVScene.

You can download the demo here.


Please keep in mind this is basically a one-man production. I'm quite satisfied with how the final demo turned out, given the limited amount of time I had to work on it.

My NVidia GTC / NVScene slides and session video

Here are the slides and video of my presentation mentioned in the previous post:

Slides (PDF)
Video (ustream)

Unfortunately, it seems something went wrong during recording, the first 6:28 are audio-only, but then video also appears..

Monday, March 17, 2014

Speaking at NVidia GPU Tech Conference / NVScene 2014

I'll give a talk at NVScene in San Francisco on Thursday 27th:

Demo Engine Tricks of the Trade (2:00 pm, Session S4819)


The presentation will cover tips to improve design workflows, increase rendering performance and how to simplify graphics engine code.

Friday, October 18, 2013

Fixing RenderMonkey Installer Fatal Error

Here is a workaround if your installation of RenderMonkey fails with "Fatal Error: Installation ended prematurely":
Open a command prompt, go to the directory where the RenderMonkey .msi file is located and type:

msiexec /qr /i RenderMonkey.2008-12-17-v1.82.322.msi

Thursday, October 10, 2013

FBX SDK Tips

Here are some little tips based on my current experience with the FBX SDK (2014.1):

Linker problems

If you get linker errors (e.g. related to ClassId when using the GetSrcObject or FindMember functions) make sure you link to "libfbxsdk-md.lib" instead of "libfbxsdk.lib".

Node Transformations

Regarding computation of node transformation matrices you need to realize that there are two different formulas used depending on whether the current file is a Maya scene or comes from 3dsmax.
Make sure you read the relevant section in the FBX SDK Programmer's Guide -> Nodes and the Scene Graph -> FBX Nodes -> "Computing Transformation Matrices" (http://docs.autodesk.com/FBX/2014/ENU/FBX-SDK-Documentation/).


The first step is to get the local node transformation. There are several ways to do this:

Option 1:
FbxMatrix matTransform = pNode->GetScene()->GetEvaluator()->GetNodeLocalTransform(pNode);

Option 2:
FbxVector4 translation = pNode->EvaluateLocalTranslation();
FbxVector4 rotation = pNode->EvaluateLocalRotation();
FbxVector4 scale = pNode->EvaluateLocalScaling();
 
FbxMatrix matTransform( translation, rotation, scale );
   
Option 3:
FbxVector4 translation = pNode->LclTranslation.Get();
FbxVector4 rotation    = pNode->LclRotation.Get();
FbxVector4 scale       = pNode->LclScaling.Get();   
FbxMatrix matTransform( translation, rotation, scale );


The next step is to get the geometric transform (assuming 3dsmax is used):

FbxVector4 T = pNode->GetGeometricTranslation (FbxNode::eSourcePivot);
FbxVector4 R = pNode->GetGeometricRotation(FbxNode::eSourcePivot);
FbxVector4 S = pNode->GetGeometricScaling(FbxNode::eSourcePivot);
FbxMatrix matGeometricOffset(T, R, S);


Then we apply the mentioned formula for 3dsmax:

FbxMatrix nodeMatrix = parentMatrix * matTransform;
FbxMatrix worldMatrix = nodeMatrix * matGeometricOffset;


The resulting matrix should then be converted to the conventions of the 3D API you are using (GL/D3D/etc.)

Getting all AnimStacks of the current Scene

The FBX SDK in many cases provides several different ways to retrieve the same information, which can be a bit confusing at times.
To iterate over all AnimStacks in the scene you can use

    for(int i=0; i < m_pImporter->GetAnimStackCount(); i++)
    {

        FbxAnimStack* pAnimStack =  m_pScene->GetSrcObject(i);
        // ...
    }

Or alternatively:

    for(int i=0; i < m_pImporter->GetAnimStackCount(); i++)
    {
        FbxTakeInfo* takeInfo = m_pImporter->GetTakeInfo(i);
        FbxString takeName = takeInfo->mName;


        FbxAnimStack* pAnimStack = m_pScene->FindMember((const char*)takeName);
        // ...
    }

But as said, there are even more ways of doing the same thing...

Friday, July 13, 2012

Open Asset Import Library (AssImp) is great!

I can only recommend you to try it out, it has a very simple and well structured API and supports many different formats, including COLLADA.
You don't even need to compile the whole thing, just download the package from sourceforge, add the include path and dynamically link to the provided DLL (you will only need to import aiImportFile, aiReleaseImport and some of the material getter functions).
I use AssImp in my little converter tool to translate COLLADA files to a custom binary format.

Here is a small snippet to dump simple meshes (separate binary files for vertex attributes and indices):

(some helper code first)
class ExportFile {
public:
    FILE* f;
 
    ExportFile(const char* filename) { f = fopen(filename, "wb"); }
    ~ExportFile() { fclose(f); }
    inline void writeU32(uint32 value) { fwrite(&value, 1, sizeof(uint32), f); }
    inline void write(void* data, size_t length) { fwrite(data, 1, length, f); }
 // etc.
};

void writeArray(const char* filename, unsigned int count, size_t stride, void* data)
{
    ExportFile f(filename);
    f.writeU32(count);
    f.write(data, count*stride);
}
Here's the export function:
void exportMesh(aiMesh* mesh)
{
    std::string meshName = meshNodeNames[mesh];

 // vertex arrays
    writeArray((meshName + std::string(".vtx")).c_str(), mesh->mNumVertices,
               sizeof(aiVector3D), mesh->mVertices);
    writeArray((meshName + std::string(".nrm")).c_str(), mesh->mNumVertices, 
                sizeof(aiVector3D), mesh->mNormals);

    if (mesh->mTangents) {
        writeArray((meshName + std::string(".tan")).c_str(), mesh->mNumVertices, 
                   sizeof(aiVector3D), mesh->mTangents);
    }
 
    // texcoords
    char uvFilename[256];
    int numUVs = 0;
    for (unsigned int i = 0; i < AI_MAX_NUMBER_OF_TEXTURECOORDS; i++)
    {
        if (mesh->mTextureCoords[i]) {
            numUVs++;
            unsigned int numComps = mesh->mNumUVComponents[i];
            aiVector3D* uvw_array = mesh->mTextureCoords[i];
            float* tex = new float[numComps * mesh->mNumVertices];

            for (unsigned int v=0; v < mesh->mNumVertices; v++) {
                if (numComps >= 1) tex[v*numComps+0] = uvw_array[v].x;
                if (numComps >= 2) tex[v*numComps+1] = uvw_array[v].y;
                if (numComps >= 3) tex[v*numComps+2] = uvw_array[v].z;
            }

            sprintf(uvFilename, "%s.uv%d", meshName.c_str(), i);
            writeArray(uvFilename, mesh->mNumVertices, sizeof(float)*numComps, tex);

            delete[] tex;
        }
    }

    // indices
    size_t idxCount = mesh->mNumFaces*3;
    assert(idxCount <= 65536);

    uint16* indices = new uint16[idxCount];
    for (unsigned int i=0; i < mesh->mNumFaces; i++) {
        assert(mesh->mFaces[i].mNumIndices == 3);
        indices[i*3+0] = mesh->mFaces[i].mIndices[0];
        indices[i*3+1] = mesh->mFaces[i].mIndices[1];
        indices[i*3+2] = mesh->mFaces[i].mIndices[2];
    }
    writeArray((meshName + std::string(".idx")).c_str(), idxCount, sizeof(uint16), indices);

    delete[] indices;
}

Tuesday, June 26, 2012

Using the D3DCompiler DLL with older DX9 SDKs via dynamic linking

If you want to compile shaders at runtime without D3DX, you may want to use the D3DCompiler_xx.dll.

Here is an example of how to dynamically load the D3DCompiler_xx.dll under DX9 without requiring a newer DXSDK installation which includes the complete D3D10/11 headers/libraries.

This approach has the added benefit that it works with the mingw compiler / QtCreator as well.

First we declare some required COM interfaces, D3D structures and the D3DCompile() function prototype (ofcourse you can add more functions like D3DCompileFromFile, etc. if needed):

#ifndef __ID3D10Blob_FWD_DEFINED__
#define __ID3D10Blob_FWD_DEFINED__
typedef interface ID3D10Blob ID3D10Blob;
#endif

#define D3DCOMPILER_DLL_A "d3dcompiler_43.dll"

typedef struct {
    LPCSTR Name;
    LPCSTR Definition;
} D3D_SHADER_MACRO;

DEFINE_GUID(IID_ID3D10Blob, 0x8ba5fb08, 0x5195, 0x40e2, 0xac, 0x58,
 0xd, 0x98, 0x9c, 0x3a, 0x1, 0x2);

typedef struct ID3D10BlobVtbl {
    BEGIN_INTERFACE

    HRESULT ( STDMETHODCALLTYPE *QueryInterface )
            (ID3D10Blob * This, REFIID riid, void **ppvObject);

    ULONG   ( STDMETHODCALLTYPE *AddRef )(ID3D10Blob * This);
    ULONG   ( STDMETHODCALLTYPE *Release )(ID3D10Blob * This);
    LPVOID  ( STDMETHODCALLTYPE *GetBufferPointer )(ID3D10Blob * This);
    SIZE_T  ( STDMETHODCALLTYPE *GetBufferSize )(ID3D10Blob * This);

    END_INTERFACE
} ID3D10BlobVtbl;

#define ID3D10Blob_QueryInterface(This,riid,ppvObject) \
    ( (This)->lpVtbl -> QueryInterface(This,riid,ppvObject) )

#define ID3D10Blob_AddRef(This) \
    ( (This)->lpVtbl -> AddRef(This) )

#define ID3D10Blob_Release(This) \
    ( (This)->lpVtbl -> Release(This) )

#define ID3D10Blob_GetBufferPointer(This) \
    ( (This)->lpVtbl -> GetBufferPointer(This) )

#define ID3D10Blob_GetBufferSize(This) \
    ( (This)->lpVtbl -> GetBufferSize(This) )

interface ID3D10Blob
{
    CONST_VTBL struct ID3D10BlobVtbl *lpVtbl;
};

typedef ID3D10Blob ID3DBlob;

typedef HRESULT (WINAPI *pD3DCompile)
    (LPCVOID                         pSrcData,
     SIZE_T                          SrcDataSize,
     LPCSTR                          pFileName,
     CONST D3D_SHADER_MACRO*         pDefines,
     void*/*ID3DInclude* */          pInclude,
     LPCSTR                          pEntrypoint,
     LPCSTR                          pTarget,
     UINT                            Flags1,
     UINT                            Flags2,
     ID3DBlob**                      ppCode,
     ID3DBlob**                      ppErrorMsgs);
Now we can load the D3DCompiler DLL, import the function via GetProcAddress and then compile our shader:
    HINSTANCE hD3DCompiler = LoadLibraryA(D3DCOMPILER_DLL_A);
    pD3DCompile D3DCompile = NULL;

    if (hD3DCompiler) {
        D3DCompile = (pD3DCompile) GetProcAddress(hD3DCompiler, "D3DCompile");
    }

    ID3DBlob* shaderBlob = NULL;
    ID3DBlob* errorMsg = NULL;
    D3DCompile(vsSource, strlen(vsSource), NULL, NULL, NULL, "main", "vs_2_0",
               0, 0, &shaderBlob, &errorMsg);
    if (errorMsg) {
        char text[1024];
        sprintf(text, "D3DCompile failed:\n%s",
               (char*) ID3D10Blob_GetBufferPointer(errorMsg));
        MessageBoxA(NULL, text, "Error", MB_OK|MB_ICONERROR|MB_TASKMODAL);
        ID3D10Blob_Release(errorMsg);
        exit(-1);
    }

    IDirect3DVertexShader9* shader = NULL;
    if (shaderBlob) {
        gD3DDevice->CreateVertexShader(
           (DWORD *) ID3D10Blob_GetBufferPointer(shaderBlob),
            &shader
        );
        ID3D10Blob_Release(shaderBlob);
    }

Thursday, November 17, 2011

Little COM smart pointer tip

Many people use CComPtr or Boost smart pointers to help with reference counting of COM/Direct3D objects, like vertex/index buffers, textures, etc. (see here for a good overview on the topic).

However CComPtr requires ATL and doesn't work with VS Express.

An alternative is to use _com_ptr_t which also is available in VS Express, with a little macro help you can wrap things up quite nicely:

#include "comip.h"

#define DECLARE_COM_REF(x) typedef \ 
_com_ptr_t< _com_IIID< IDirect3D##x##9, &IID_IDirect3D##x##9 > > \
x##Ref;

DECLARE_COM_REF(VertexBuffer);
DECLARE_COM_REF(IndexBuffer);
DECLARE_COM_REF(CubeTexture);
// etc.

With this you can then create some nice wrapper classes like this e.g.:

class VertexBuffer : public VertexBufferRef {
public:
    VertexBuffer(IDirect3DVertexBuffer9* vertex_buffer = NULL) { 
        Attach(vertex_buffer);
    }

    void* lock(UINT offset, UINT size, DWORD flags)
    {
        IDirect3DVertexBuffer9* p = GetInterfacePtr();
        assert(NULL != p);
        void* mem = 0;
        p->Lock(offset, size, &mem, flags);
        assert(NULL != mem);
        return mem;
    }

    void unlock()
    {
        GetInterfacePtr()->Unlock();
    }
};

Wednesday, January 12, 2011

Deferred Rendering: Reconstructing Position from Depth

Here is a short snippet of how to reconstruct view-/world-space position from depth in deferred rendering.

This was originally described by fpuig, but here is the cleaned up and bugfixed (hopefully..) version.

What's nice about this code is that it works both for light volume geometry (e.g. spheres for omni lights) and fullscreen quads:

1) In your GBuffer pass store positionInViewSpace.z in the depth rendertarget.

2) The lighting vertex-shader calculates the eye-to-pixel rays (in view-/world-space):
    OUT.position =  mul(matrixWVP, float4(IN.position,1));
    OUT.vPos = ConvertToVPos(OUT.position); // sm2 has no VPOS..
    OUT.vEyeRayVS = float3(OUT.position.x*TanHalfFOV*ViewAspect, 
                           OUT.position.y*TanHalfFOV, OUT.position.w);
    OUT.vEyeRay = mul(matrixViewInv, OUT.vEyeRayVS);
In shader model 2 we don't have the VPOS interpolator, so we can use the following function (RTWidth/Height is the size of your screen/rendertargets):
    float4 ConvertToVPos(float4 p)
    {
        return float4(0.5*(float2(p.x+p.w, p.w-p.y) + 
                      p.w*float2(1.0f/RTWidth, 1.0f/RTHeight)), p.zw);
    }
3) Then, in the pixel-shader one can compute the view-space and world-space position per pixel:
    float depth = tex2Dproj(GBufferDepthSampler, IN.vPos); // divide by W
    IN.vEyeRayVS.xyz /= IN.vEyeRayVS.z; // divide by W
    float3 pixelPosVS = IN.vEyeRayVS.xyz * depth;
    float3 pixelPosWS = IN.vEyeRay.xyz * depth + cameraPosition.xyz;

This can be optimized a lot, of course.

Thursday, September 16, 2010

Optimizing Vertex Formats

32 bytes per vertex is optimal for the hardware vertex cache.

We can safely pack normals and tangents to a uint32 each.

Also, I mainly use the second UV set with unique mappings / lightmaps, so it is ensured that the coords always lie within [0,1] - this allows to use D3DDECLTYPE_SHORT2N (the unnormalized SHORT2 type is not supported on my trusty old ATi x700 mobility..). To convert your float UVs you just multiply them with 32767.0f

As a result we get a nice vertexformat like this:

// size = 32 bytes
struct Vertex {
 FVec3 pos;
 VecU32 nrm;
 FVec2 uv;
        short uv2[2];
 VecU32 tan;
};

D3DVERTEXELEMENT9 declExt[] = {
 // stream, offset, type, method, usage, usageIndex
 { 0, 0, D3DDECLTYPE_FLOAT3, D3DDECLMETHOD_DEFAULT, D3DDECLUSAGE_POSITION, 0 },
 { 0, 12, D3DDECLTYPE_UBYTE4, D3DDECLMETHOD_DEFAULT, D3DDECLUSAGE_NORMAL, 0 },
 // 2d uv
 { 0, 16, D3DDECLTYPE_FLOAT2, D3DDECLMETHOD_DEFAULT, D3DDECLUSAGE_TEXCOORD, 0 },
 { 0, 24, D3DDECLTYPE_SHORT2N, D3DDECLMETHOD_DEFAULT, D3DDECLUSAGE_TEXCOORD, 1 },
 // tangent
 { 0, 28, D3DDECLTYPE_UBYTE4, D3DDECLMETHOD_DEFAULT, D3DDECLUSAGE_TEXCOORD, 2 },
 D3DDECL_END()
};

In your shader just use:

struct VertexInput {
 float4 position : POSITION;
 float4 norm : NORMAL;  // compressed uint32
 float2 uv : TEXCOORD0;
 float2 uv2 : TEXCOORD1;
 float4 tangent : TEXCOORD2; // compressed uint32
};

VertexOutput main(VertexInput IN) {
 ...
 // decompress normal & tangent
 float3 N = 2.0f*IN.norm/255.0f-1.0f;
 float4 T = 2.0f*IN.tangent/255.0f-1.0f;
 ...
}

Here is some C++ code to de-/compress your normals:

class VecU32 {
public:
 union {
  u8 dir[4];
  u32 vec32;
 };

 inline void compress(const FVec3& nrm) {
  dir[0] = (u8)((nrm.x * 0.5f + 0.5f) * 255.0f);
  dir[1] = (u8)((nrm.y * 0.5f + 0.5f) * 255.0f);
  dir[2] = (u8)((nrm.z * 0.5f + 0.5f) * 255.0f);
  dir[3] = 255;
 }

 inline void compress(const FVec4& nrm) {
  dir[0] = (u8)((nrm.x * 0.5f + 0.5f) * 255.0f);
  dir[1] = (u8)((nrm.y * 0.5f + 0.5f) * 255.0f);
  dir[2] = (u8)((nrm.z * 0.5f + 0.5f) * 255.0f);
  dir[3] = (u8)((nrm.w * 0.5f + 0.5f) * 255.0f);
 }

 inline FVec4 decompress() {
  return FVec4(
   2.0f*dir[0]/255.0f-1.0f,
   2.0f*dir[1]/255.0f-1.0f,
   2.0f*dir[2]/255.0f-1.0f,
   2.0f*dir[3]/255.0f-1.0f
   );
 }
};

Wednesday, September 8, 2010

Light Pre-Pass Rendering

This is my implementation of Light Pre-Pass Rendering. 


It's a deferred lighting technique (similar to deferred shading, but without the need for big G-Buffers) - Basically, in a first pass you render depth and normals, then accumulate lighting values of all lights into a light buffer rendertarget (using the normals stored in the first pass) and finally in the third step you compose the lighting info with your material info (textures, etc.) during forward rendering.

You should check the presentations by Wolfgang Engel and others for more background info, the method is currently very popular on PS3/XBox (Uncharted, Resistance2, LBP, Blur, to name a few).

Some notes on the code:
  • the ZN pre-pass stores depth and view-space normals packed in a single ARGB8888 rendertarget using 2x8-bit components each
  • to encode the normals I use the old N.z = sqrt(1-dot(N.xy, N.xy) trick which is not 1oo% correct, but works OK for now
  • to en/decode the depth values I came up with the following formulas, didn't check yet if there are better solutions (Note: farZ = cameraFarPlane / 256.0) :
float2 EncodeFloat16(float v) { 
    float fac = v / 256.0f; 
    float fra = frac(fac); 
    return float2((fac-fra) / farZ, fra); 
} 

float DecodeFloat16(float2 v) { 
    return (v.x * farZ + v.y) * 256.0f; 
} 
  • compositing the lighting info in the second (forward) geometry pass is done in linear color space: textures are converted from gamma2 to gamma1 during texture fetch, and the final result is converted back to gamma2 for display
  • the ZN ARGB8888 packing of course causes some artifacts, but it's still acceptable in my current test scenes

Monday, August 30, 2010

Books

Two books I enjoyed reading recently are:

Programming with POSIX Threads by David R. Butenhof (Addison-Wesley)

Very well written, concise and nicely layouted. Additionally to discussing the posix functions in detail, the second half of the book contains further material regarding worker thread pools, synchronization techniques, etc. Highly recommended.

Hacker's Delight by Henry S. Warren (Addison-Wesley, 2003)

Bit-twiddling to the max! This came in handy several times already, when needing to squeeze out some additional CPU cycles in innerloops.
See also: http://www.hackersdelight.org/

Check them out, you might also like them...

Tuesday, August 24, 2010

Assembly 2010

One of the best examples of creativity and skilled programming of GPU effects could be found at Assembly this year, you should definitely check out the released demos and intros.

ASD's "Happiness is around the bend" made 1st place in the competition:

Gamescom

My personal Gamescom highlight was clearly Enslaved by Ninja Theory. I just love the scenario, atmosphere and graphics! I hope the final gameplay will also be improved over Heavenly Sword, but let's see..

Here are some links and pictures:
http://uk.ps3.ign.com/articles/107/1079980p1.html
http://www.joystiq.com/tag/enslaved

Second highlight for sure were the FMVs of Star Wars The Old Republic - wow, they seemed actually higher quality than the films' CGs (except for the super-obvious matte paintings in the background during the first minutes)..
http://www.swtor.com/media/trailers/hope-cinematic-trailer

But the game itself is not my cup of tea, though, I fear...

The rest of the show? Lost in the blurrines of decibels, Kinect hype, 2h waiting queues and hordes of people..

Wednesday, August 4, 2010

Ambient occlusion tests


Testing various features of the engine/3dsmax exporter: multiple uv-channels, ambient occlusion texture baking, etc.