Maruf Notes

Thursday, July 15, 2010

Just In Time Compiler for Managed Platform- Part 2: Generate native method

Today we'll design a small block of code that is equivallent to a corresponding java method.

Since the generated native executable code will be used only by our VM we are free to define out own structure and calling convention for it. We generate one native function for each Java method. Each generated native function will have only parameter that is required for operation- a pointer to a structure RuntimeEnvironment:

union Variable
{
    u1 charValue;
    u2 shortValue;
    u4 intValue;
    f4 floatValue;
    u4* ptrValue;
    Object object;
};

struct RuntimeEnvironment
{
    Variable *stack;
    int stackTop;
    //We'll add more as we require later
};

The return type will be int:

int ExecuteMethod(RuntimeEnvironment *pRE);

To generate the final method we need a lot of helper function. We define those as:

void HelperFunction(u1* code, int& ip);

These functions will take a code block and insert code in the block and fix the code pointer (ip).

First let us define the Prolog and Epilog and return 0 helper functions for this function prototype:

void Prolog(u1* code, int& ip)
{
    u1 c[]= {
        0x55,//               push        ebp  
        0x8B, 0xEC,//            mov         ebp,esp 
        0x81, 0xEC, 0xC0, 0x00, 0x00, 0x00,// sub         esp,0C0h 
        0x53,//               push        ebx  
        0x56,//               push        esi  
        0x57,  //             push        edi 
    };

    memcpy(&code[ip], c, sizeof(c));
    ip+=sizeof(c);
}

void Epilog(u1* code, int& ip)
{
    u1 c[]= {
        0x5F,//               pop         edi  
        0x5E,//               pop         esi  
        0x5B,//               pop         ebx  
        0x8B, 0xE5,//            mov         esp,ebp 
        0x5D,//               pop         ebp  
        0xC3,//               ret              
    };

    memcpy(&code[ip], c, sizeof(c));
    ip+=sizeof(c);
}

void Return0(u1* code, int& ip)
{
    //33 C0  xor         eax,eax 
    code[ip++]=0x33;
    code[ip++]=0xC0;
}

Now, we want to generate machine code for the following simple function-

public static int SimpleCall()
{
    return 17;
}

Here is the generated java byte code:

Signature: ()I
  Code:
    0:   bipush  17
    2:   ireturn

Here we start to generate helper function for Java Virtual Machine instruction.

For bipush [value] we need to push the [value] on the VM stack:

void BiPush(u1* code, int& ip, short value)
{
    // C++ equivallent
    // pRE->stack[pRE->stackTop++].shortValue = value;

    u1 c[] = {
         0x8B, 0x45, 0x08, //         mov         eax,dword ptr [pRE] 
         0x8B, 0x48, 0x04, //         mov         ecx,dword ptr [eax+4] 
         0x8B, 0x55, 0x08, //         mov         edx,dword ptr [pRE] 
         0x8B, 0x02, //            mov         eax,dword ptr [edx] 
         0xBA, 0x00, 0x00, 0x00, 0x00, //   mov         edx,value 
         0x66, 0x89, 0x14, 0xC8, //      mov         word ptr [eax+ecx*8],dx 
         0x8B, 0x45, 0x08, //         mov         eax,dword ptr [pRE] 
         0x8B, 0x48, 0x04, //         mov         ecx,dword ptr [eax+4] 
         0x83, 0xC1, 0x01, //         add         ecx,1 
         0x8B, 0x55, 0x08, //         mov         edx,dword ptr [pRE] 
         0x89, 0x4A, 0x04,  //       mov         dword ptr [edx+4],ecx 
    };

    //We need to encode value and set it to the 00 00 00 00 position 
    u1 encVal[4];
    EncodeByte4((int)value, encVal);
    memcpy(c + 12, encVal, 4); 
    memcpy(&code[ip], c, sizeof(c));
    ip+=sizeof(c);
}

Thats it for the simple java function. We can now test this:

int main() 
{ 
    int codeBlockSize = 4096;    
    int (*SimpleCall)(RuntimeEnvironment *)=(int (*)(RuntimeEnvironment *)) VirtualAlloc(NULL, codeBlockSize,  MEM_COMMIT, PAGE_EXECUTE_READWRITE);
    u1* codes = (u1*) SimpleCall;
    int ip=0;
    memset(codes, 0, codeBlockSize);

    Prolog(codes, ip);
    BiPush(codes, ip, 17);
    Return0(codes, ip);
    Epilog(codes, ip);

    //No lets test if it is really pushing value 17 on the VM stack
    RuntimeEnvironment *pRE = new RuntimeEnvironment();;
    pRE->stack = new Variable[20];    
    memset(pRE->stack, 0, sizeof(Variable)*20);
    int retVal = (*SimpleCall)(pRE);
    printf("pRE->stack[0].intValue = %d", pRE->stack[0].intValue);    
    
    return 0;
}

Thats cool- we have generated our first native function that actually does some byte code execution.

Saturday, July 10, 2010

Just In Time Compiler for Managed Platform- Part 1: Code Generation

First we need to generate executable code block.

So let us write some code in C++.

int add(int x, int y) 
{ 
    int r; 
    r= x+y; 
    return r; 
}

int main()
{
    int r1 = add(13, 23);
    printf("Returned value = %d", r1);
}

Great! It returns the right value. OK, but that is very basic. We want to generate the function from data in a simple buffer-

First we need a machine equivallent code for the function above:

unsigned char addcode[] = {  
    0x55,          //push        ebp  
    0x8B, 0xEC,        //mov         ebp,esp 
    0x81, 0xEC, 0xC0, 0x00, 0x00, 0x00,  //sub         esp,0C0h 
    0x53,          //push        ebx  
    0x56,          //push        esi  
    0x57,          //push        edi  

    //r=x+y;
    0x8B, 0x45, 0x08,       //mov         eax,dword ptr [x] 
    0x03, 0x45, 0x0C,       //add         eax,dword ptr [y] 
    0x89, 0x45, 0xF8,       //mov         dword ptr [r],eax 

    //return r;
    0x8B, 0x45, 0xF8,       //mov         eax,dword ptr [r] 

    0x5F,          //pop         edi  
    0x5E,          //pop         esi  
    0x5B,          //pop         ebx  
    0x8B, 0xE5,        //mov         esp,ebp 
    0x5D,          //pop         ebp  

    0xC3          //ret              
   };

OK, this code is generated by Visual Studio compiler. We use it to generate our own code block in memory.

To get a memory block we can use to generate executable code we use the following Windows API:

LPVOID WINAPI VirtualAlloc(
    __in_opt LPVOID lpAddress,
    __in SIZE_T dwSize,
    __in DWORD flAllocationType,
    __in DWORD flProtect
);

First we allocate a 4096 byte executable code block and put it in a function pointer:

int (*addfn)(int, int) = (int (*)(int, int)) VirtualAlloc(NULL, 4096,  MEM_COMMIT, PAGE_EXECUTE_READWRITE);

Then we copy our executable code to this memory block:

memcpy(addfn, addcode, sizeof(addcode));

Now the majic - we call the function and print the return value:

int r1 = (*addfn)(13,23); 
 printf("Returned value = %d", r1);

Thats easy- right?

Now we release the memory since we are gentle citizen-

VirtualFree(addfn, NULL, MEM_RELEASE);

Thats it. here is the full code:

#include [windows.h][stdio.h]...

unsigned char addcode[] = {  
    0x55,          //push        ebp  
    0x8B, 0xEC,        //mov         ebp,esp 
    0x81, 0xEC, 0xC0, 0x00, 0x00, 0x00,  //sub         esp,0C0h 
    0x53,          //push        ebx  
    0x56,          //push        esi  
    0x57,          //push        edi  

    //r=x+y;
    0x8B, 0x45, 0x08,       //mov         eax,dword ptr [x] 
    0x03, 0x45, 0x0C,       //add         eax,dword ptr [y] 
    0x89, 0x45, 0xF8,       //mov         dword ptr [r],eax 

    //return r;
    0x8B, 0x45, 0xF8,       //mov         eax,dword ptr [r] 

    0x5F,          //pop         edi  
    0x5E,          //pop         esi  
    0x5B,          //pop         ebx  
    0x8B, 0xE5,        //mov         esp,ebp 
    0x5D,          //pop         ebp  

    0xC3          //ret              
   };

int main()
{
    int (*addfn)(int, int) = (int (*)(int, int)) VirtualAlloc(NULL, 4096,  MEM_COMMIT, PAGE_EXECUTE_READWRITE);
    memcpy(addfn, addcode, sizeof(addcode)); 
    
    int r1 = (*addfn)(13,23); 
    printf("Returned value = %d", r1);

    VirtualFree(addfn, NULL, MEM_RELEASE);

    return 0;
}

Thats all for now. We can now generate code in memory and execute it. First step to design a JIT compiler.

Wednesday, April 30, 2008

Compiler in action- C/C++ to Machine

Introduction

What happens when I give my C/C++ code to a compiler? It generates machine code.
But I want to know what machine code it generates really. I use the compiler that comes Visual C++ 2008.
Other versions should be similar if not same.

Producing Assembly Output

With Visual Studio we can produce assembly language output with following settings:

Project Property Pages > Configuration Properties > C++ > Output Files

Assembler Output: Assembly With Source Code (/FAs)

The compiler generates assembly code and output with corresponding C/C++ source code. Its very
useful to understand how the compiler works.

Function

A function when compiled has its prolog, epilog and ret instructions along with its body.
It maintains the stack and local variables.

Prolog and Epilog

Prolog is a set of instructions that compiler generates at the beginning of a function and epilog is
generated at the end of a function. This two maintains stack, local variables, registers and unwind
information.

Every function that allocates stack space, calls other functions, saves nonvolatile registers, or
uses exception handling must have a prolog whose address limits are described in the unwind data
associated with the respective function table entry. The prolog saves argument registers in their home
addresses if required, pushes nonvolatile registers on the stack, allocates the fixed part of the stack
for locals and temporaries, and optionally establishes a frame pointer. The associated unwind data must
describe the action of the prolog and must provide the information necessary to undo the effect of the
prolog code [MSDN].

Let us see what is generated as prolog and epilog. We have a function named add like this:

int add(int x, int y)
{
 int p=x+y;
 return p;
}

And the generated assembly listing:

_p$ = -4      ; size = 4
_x$ = 8       ; size = 4
_y$ = 12      ; size = 4
?add@@YAHHH@Z PROC     ; add, COMDAT
; 12   : {
;Prolog
push ebp
mov ebp, esp
push ecx

; 13   :  int p=x+y;
mov eax, DWORD PTR _x$[ebp]
add eax, DWORD PTR _y$[ebp]
mov DWORD PTR _p$[ebp], eax
; 14   :  return p;
mov eax, DWORD PTR _p$[ebp]
; 15   : }

;Epilog
mov esp, ebp
pop ebp

ret 0 ;disposition of stack- 0 disp as returning through register
?add@@YAHHH@Z ENDP     ; add

Not much work. The compiler just saves EBP register copies the ESP register in EBP register and use
EBP as stack pointer at prolog and at epilog stage it restores the EBP register. Sometimes there is a
subtraction to handle local variables. There are two instruction ENTER and LEAVE that can be used
in place of push pop things.

Function Parameters/Local Variables

The function parameters are placed at positive offset from the stack pointer and local variables are
located at negative offset at the time of calling the function.
Function parameters are pushed on the stack before calling and the function may initialize the
local variable. From previous assembly listing we find parameters x and y is at offset 8 and 12
and the local variable p is at offset -4 from the stack top.

Function call

The CALL instruction is used to invoke a function. Before doing so the caller function pushes
parameter values or set register (this pointer) and issue CALL instruction. After returning the
caller function may need to set stack pointer depending on calling convention it used. We discus
this in next subsection.

Calling conventions

There are several calling conventions. Calling convention tells compiler how the parameters
are passed, how stack is maintained and how to decorate the function names in object files.
Following table shows basic things at a glance:

Calling Convention	Argument Passing	Stack Maintenance	Name Decoration (C only)	Notes
__cdecl	Right to left.	Calling function pops arguments from the stack.	Underscore prefixed to function names. Ex: _Foo.
__stdcall	Right to left.	Called function pops its own arguments from the stack.	Underscore prefixed to function name, @ appended followed by the number of decimal bytes in the argument list. Ex: _Foo@10.
__fastcall	First two DWORD arguments are passed in ECX and EDX, the rest are passed right to left.	Called function pops its own arguments from the stack.	A @ is prefixed to the name, @ appended followed by the number of decimal bytes in the argument list. Ex: @Foo@10.	Only applies to Intel CPUs. This is the default calling convention for Borland compilers.
thiscall	this pointer put in ECX, arguments passed right to left.	Calling function pops arguments from the stack.	None.	Used automatically by C++ code.
naked	Right to left.	Calling function pops arguments from the stack.	None.	Only used by VxDs.

Source: Debugging Applications by John Robbins

StdCall


int _stdcall StdCallFunction(int x, int y)
{
 return x;
}

The generated code is like this:


_x$ = 8       ; size = 4
_y$ = 12      ; size = 4
?StdCallFunction@@YGHHH@Z PROC    ; StdCallFunction, COMDAT
; 7    : {
push ebp
mov ebp, esp
; 8    :  return x;
mov eax, DWORD PTR _x$[ebp]
; 9    : }
pop ebp
ret 8
?StdCallFunction@@YGHHH@Z ENDP    ; StdCallFunction

To call the compiler generates code like this:


; 26   :  r=StdCallFunction(p, q);

mov eax, DWORD PTR _q$[ebp]
push eax
mov ecx, DWORD PTR _p$[ebp]
push ecx
call ?StdCallFunction@@YGHHH@Z ; StdCallFunction
mov DWORD PTR _r$[ebp], eax

Cdecl

The function declaration uses _cdecl keyword.


int _cdecl CDeclCallFunction(int x, int y)
{
 return x;
}

Compiler generates following assembly listing:


_x$ = 8       ; size = 4
_y$ = 12      ; size = 4
?CDeclCallFunction@@YAHHH@Z PROC   ; CDeclCallFunction, COMDAT

; 12   : {

push ebp
mov ebp, esp

; 13   :  return x;

mov eax, DWORD PTR _x$[ebp]

; 14   : }

pop ebp
ret 0
?CDeclCallFunction@@YAHHH@Z ENDP   ; CDeclCallFunction

To call the function compiler generates following code:


; 27   :  r=CDeclCallFunction(p, q);

mov edx, DWORD PTR _q$[ebp]
push edx
mov eax, DWORD PTR _p$[ebp]
push eax
call ?CDeclCallFunction@@YAHHH@Z  ; CDeclCallFunction
add esp, 8
mov DWORD PTR _r$[ebp], eax

Fastcall


int _fastcall FastCallFunction(int x, int y)
{
return x;
}

The generated code:


_y$ = -8      ; size = 4
_x$ = -4      ; size = 4
?FastCallFunction@@YIHHH@Z PROC    ; FastCallFunction, COMDAT
; _x$ = ecx
; _y$ = edx
; 17   : {
push ebp
mov ebp, esp
sub esp, 8
mov DWORD PTR _y$[ebp], edx
mov DWORD PTR _x$[ebp], ecx
; 18   :  return x;
mov eax, DWORD PTR _x$[ebp]
; 19   : }
mov esp, ebp
pop ebp
ret 0
?FastCallFunction@@YIHHH@Z ENDP    ; FastCallFunction

And to call the function:


; 28   :  r=FastCallFunction(p, q);

mov edx, DWORD PTR _q$[ebp]
mov ecx, DWORD PTR _p$[ebp]
call ?FastCallFunction@@YIHHH@Z  ; FastCallFunction
mov DWORD PTR _r$[ebp], eax

Thiscall

Used for class member functions. We discuss it later in detail.

Nacked

This calling conven is used for VxD drivers.

Representation of a class

A class is just a structure of varuables with functions. While creating an object compiler reserves
space on heap and call the constructure of the class. A class can have a table of functions (the vtable)
as the first member. It is used to call virtual functions. Class member functions are treated similar
as normal C functions with the exception that it receives this pointer as one parameter in the
ECX register.

Class Member Functions

Here is a simple class for demonstration of member function.


class Number
{
 int m_nMember;
 public:
 void SetNumber(int num, int base)
 {
  m_nMember = num;
 }
};

The SetNumber in class Number generates following listing:


_this$ = -4      ; size = 4
_num$ = 8      ; size = 4
_base$ = 12      ; size = 4
?SetNumber@Number@@QAEXHH@Z PROC   ; Number::SetNumber, COMDAT
; _this$ = ecx

; 30   :  {

push ebp
mov ebp, esp
push ecx
mov DWORD PTR _this$[ebp], ecx

; 31   :   m_nMember = num;

mov eax, DWORD PTR _this$[ebp]
mov ecx, DWORD PTR _num$[ebp]
mov DWORD PTR [eax], ecx

; 32   :  }

mov esp, ebp
pop ebp
ret 8
?SetNumber@Number@@QAEXHH@Z ENDP   ; Number::SetNumber

Call member function SetNumber of the class. The thiscall convension is used- this
parameter is passed in ECX register:


; 42   :  Number nObject;
; 43   :  nObject.SetNumber(r, p);

mov ecx, DWORD PTR _p$[ebp]
push ecx
mov edx, DWORD PTR _r$[ebp]
push edx
lea ecx, DWORD PTR _nObject$[ebp]
call ?SetNumber@Number@@QAEXHH@Z  ; Number::SetNumber

Virtual Functions

In case of virtual functions the compiler does not call a function of a classe directly. It rather
maintains table (called vtable) of function pointer for each class and while creating object of
a class assigns the corresponding classes vtable as the first member of the class. The function
call is indirect through this tables entry.

Let us create two classes with virtual functions here.


//A class with 2 virtual functions
class VirtualClass
{
public:
  VirtualClass()
 {
 }
 virtual int TheVirtualFunction()
 {
   return 1;
 }
 virtual int TheVirtualFunction2()
 {
  return 2;
 }
};


//Subclass
class SubVirtualClass: public VirtualClass
{
public:
 SubVirtualClass()
 { 
 }

 virtual int TheVirtualFunction()
 {
  return 3;
 }
};

Here is vtable of class VirtualClass.


CONST SEGMENT
??_7VirtualClass@@6B@ DD FLAT:??_R4VirtualClass@@6B@ ; VirtualClass::`vftable'
 DD FLAT:?TheVirtualFunction@VirtualClass@@UAEHXZ
 DD FLAT:?TheVirtualFunction2@VirtualClass@@UAEHXZ
 DD FLAT:__purecall
CONST ENDS

Please note that we have a table of three entry with one entry set to NULL (__purecall). This will
be assigned in subclass. Without this pure virtual function in source class we could create an object
and call the two virtual functions that would be base classes.

And the SubNumber classe's vtable is like this:


CONST SEGMENT
??_7SubVirtualClass@@6B@ DD FLAT:??_R4SubVirtualClass@@6B@ ; SubVirtualClass::`vftable'
 DD FLAT:?TheVirtualFunction@SubVirtualClass@@UAEHXZ
 DD FLAT:?TheVirtualFunction2@VirtualClass@@UAEHXZ
 DD FLAT:?PureVirtualFunction@SubVirtualClass@@UAEHXZ
CONST ENDS

We get all three virtual functions assigned here. As we did not override the
TheVirtualFunction2 function we have the base classes pointer in the subclasses
vtable- expected.

OK, but we must set the table as first member of a class object, right? Its done in the constructor.
Here is the constructor of subclass:


; 64   :  SubVirtualClass()
 push ebp
 mov ebp, esp
 push ecx
 mov DWORD PTR _this$[ebp], ecx
 mov ecx, DWORD PTR _this$[ebp]     ;we get this pointer
 ;lets call base classes constructor here
 call ??0VirtualClass@@QAE@XZ   ; VirtualClass::VirtualClass
 mov eax, DWORD PTR _this$[ebp]
 mov DWORD PTR [eax],OFFSET ??_7SubVirtualClass@@6B@  ;the vtavle set now

Conclusion

Thats all for now. I want to add Inheritance, Polymorphism, Operator Overloading, Event mechanism,
Template, COM Programming and exception handling in future.

Monday, April 28, 2008

Longest month of my life

April 2008. What really makes time shorter or longer? I do not think its velocity. Its waiting for something you really want but do not know if you are going to get that.

Wednesday, September 5, 2007

Switching on Silverlight

“Web is approaching the desktop” – Wahid Bhai said while demonstrating the new Flex project at KAZ. Flex is Flash but on steroids… well to be truthful Flex is Flash with libraries and an easy programming model and the nicest of all - a good eclipse based IDE.

Microsoft has of course an answer to Flex - The Silverlight which used to be known as WPF/e for a while. In Microsoft’s wording:

“Silverlight is Microsoft’s latest technology to design interactive web client with truly object oriented language C# with full support of VB .NET and other .NET citizens. Microsoft® Silverlight™ is a cross-browser, cross-platform plug-in for delivering the next generation of .NET based media experiences and rich interactive applications for the Web … Silverlight offers a flexible programming model that supports AJAX, VB, C#, Python, and Ruby, and integrates with existing Web applications. Silverlight supports video to all major browsers running on the Mac OS or Windows…. ” (the babble continues promising cure for cancer, food for Sudan and even the impossible - patriotism for the Bangladeshi middle class!)

Back on Earth, we the “Psycho group” at KAZ had a week of free time in between projects. So in the time honored way at KAZ we decided to do a spacer project to try out silverlight in all its Alpha glory and MS’s super hype! We decided to make a webtop with silverlight that would show an abstracted backend filesystem exposed by WCF.

This post is about our pains and joys during the project. Not gifted with great writing skills (apart from the coding kind) I will try putting down disperate items that we learnt or I felt like telling about our project.

First principle

Starting from Silverlight 1.1 the managed code is supported at client side. We decided to be very strict in this project to use only managed code for programming- no JavaScript. Like all principles this soon turned out to be a pain in a not so polite place

Installing Silverlight

To develop with Silverlight 1.1 you need Visual Studio 2008 (code name Orcas). While writing this downloadable beta version is available from Microsoft site. You also need Silverlight 1.1 plug-in distribution – alpha is available while writing this. Though not necessary the Silverlight 1.1 SDK is highly recommended. The SDK includes controls, samples, documents which may be useful to start. Please note that the Silverlight plug-in is required on any client from where the site is viewed.

Controls that comes from Boss

None!

Well not exactly true, but close enough. Silverlight 1.1 alpha distribution comes with a very limited set of controls. It has controls like rectangle, ellipse and label, canvas. It does not include button, edit box, scroll bar or any other advanced control. The SDK includes some controls like scrollbars.

Our Architecture

With no major controls available and no infrastructure for a shell, we soon realized that we need a petzold like approach to the project. We actually need to create a windowing system and the bare basics of message management.

Our final designe came out with 3 layers. Top layer is the application layer where user applications run. Middle layer is kernel which controls the events and communication between applications. It also provides a set of API for application developers. The lowest layer is an abstraction layer to communicate with the web server.

The system provides API’s for application developers for the platform while hiding the http calls – making the application development similar to desktop application development. Web call abstraction layer does that for user application.

The Kernel

The controller behind the scene controls form events, focus of controls, controls and windows to common file operations. The kernel has four main parts:

• Messaging System and Focus Manager
• Window Manager
• Process Manager
• File system Driver
• Resource Manager

Messaging System and Focus Manager

An event first comes to this manager and depending on the type routed to controls. If the user clicks on a control the focus manager updates the focus of controls. It keeps track of current on focus control. If it finds a change in focus it send OnLostFocus call to old control and OnSetFocus call to new focused control. With this exception most messages are routed to the focused control on arrival.

Window Manager

The window manager keeps track of each window that is created on the system. At any time window manager provides the list of windows currently available in the system. A utility application like TaskManager or TaskBar can use the list for display and control the windows. In our case the TaskBar buttons are created from the list and user can control minimize or restore operation from the TaskBar application. To keep track of the windows the window manager uses an internal generic Window list. When a new window is created the window is registered with window manager and when a window is disposed the window is unregistered.

Process Manager

The SilverlightDesktop applications that are created by implementing the IApplication interface in the main class are managed by the process manager. The process manager exposes a service to create a new process. It accepts class name as its first parameter and other parameters are passed through a Parameter object that can contain strings. On success the process manager returns a Instance object that can identify the process. Simplified version of Instance class is like this –

public class Instance{ public int Handle; public int ParentHandle; public string Name; }

The returned interface object can uniquely identify the process and it is required to handle the process.

Designing a new control

ControlBase class can be extended to design a custom control. The control may have a xaml also. The ControlBase getter ResourceName of type string should be overirden to specify the name of the xaml resource in the assembly. It is also possible to set the assembly of the control. The constructor of ControlBase class iterates through each resource and look for the specified resource to get the right xaml resource. Same technique is used in the Silverlight SDK. For example we want to have a icon control that has a image and text. Our xampl file would be similar to this:

<Canvas xmlns=”http://schemas.microsoft.com/client/2007”
xmlns:x=”http://schemas.microsoft.com/winfx/2006/xaml”
Width=”32″ Height=”50″ x:Name=”_container" >
<Canvas x:Name=”_icon” Width=”32″ Height=”32″ Canvas.Left=”0″ Canvas.Top=”0″/>

<TextBlock x:Name=”_name” Width=”80″ Height=”15″ Canvas.Left=”-24″ Canvas.Top=”32″ />
</Canvas>

We used a Canvas for icon, a TextBlock for label and a container canvas to hold them both.

And our code back for the control would be:

We used a Canvas for icon, a TextBlock for label and a container canvas to hold them both.And our code back for the control would be:

public class Icon : ControlBase{
Canvas icon;
TextBlock text;
Canvas container;
protected override void OnInitialized() {
icon = actualControl.FindName(”_icon”) as Canvas;
text = actualControl.FindName(”_name”) as TextBlock;
container = actualControl.FindName(”_container”) as Canvas;
}
protected override string ResourceName {
get {
return “Icon.xaml”;
}
}
}

Now if we add some getter/setter we get the icon control ready to be used.

The Window

Window is a customized control- but is different from others. It uses mouse events to implement feature like drag/ drop, has a title bar and sizing buttons for minimize to system taskbar, maximize to cover full user desktop area or close button to close and dispose the control from the system.

It was first posted at http://www.kaz.com.bd/blog/index.php/?p=10

Wednesday, June 6, 2007

MISL to C# (Sharp) -> Loops - the simple cases

I wish, as a decompiler writer, there would be no loops. Programmers use thousands of goto statements with if statements. But as it is not the case I must understand how to parse MSIL instructions that were generated from the loops.

Of the three types of most common loops (for, while, do-while) the while loop is the basic one. The block that is generated from any type of these loops usually has a conditional jump (usually a brtrue.s ) as the last instruction of the block. The difference from if structure is- the instruction jumps to an offset less than current instruction offset. The for and while loop has a unconditional branch (br.s ) to an offset that is between start and end of the block. The jump target usually at the beginning of the condition checking instructions. The do-while loop lacks this branch for the reason - it does not test the condition before it is at the end of the block.

So, we get instruction block like following MSIL block:

-------------------------------------------------------------------------------------
IL_0010: br.s IL_005a ;do-while loop does not have this line
IL_0012: nop

[any type and number of instructions]

IL_0059: nop

[condition check instruction- results boolean value on stack]

IL_0060: brtrue.s IL_0012
-------------------------------------------------------------------------------------

On my previous post on if structure I showed how to create a boolean condition for if structure. Things are similar here for the loops. Follow the instructions - get the top stack element when conditional jump found - reverse it (add just an !) for brtrue.s jump and put it as the loop statements condition statement. Please note that the conditional jump targets the instruction just at or after the starting instruction of the block.

Here we find we can not have a single passing decompiler. We must identify the code blocks in an iteration before final iteration. Till now we can identify blocks of if, for, while, do-while structures by using conditional jump instructions and their destination. If has destination offset after the current offset and others have destination before the current offset. The for and while can not be distinguished very clearly but the do-while does not have a jump at the beginning. And of course there can be nested blocks that are generates from nested loops.

There is some complex variation of the loops - like infinite loops, forcach loop etc. They are not much different. But I want to discuss the control structure later in more detail. This post is little bit more theoretical- see you again very soon with more interesting things.

Monday, June 4, 2007

MISL to C# (Sharp) -> Iff (if and only if)

Today I’ll talk about the most basic and useful if-then structure. The compiler generates a conditional branch. We create a block of instructions for the structure. The block is separated by the branch instruction (like brtrue) and the branch label (like IL_0019 - where the code jumps). And our condition is on the stack. If we find a true condition branch we negate it and put it as a if statement condition. The block is initially a block of MSIL that we will convert to C# code later (
may need recursion here??). Please note that the labels are not stored in MSIL. It is just the byte offset of the MSIL in a method.

If you have not already read you are requested to read my previous post.

We take a simple method to test our theory.

public int IfStructure(int a, int b)
{
bool CS_4_0001;
if (a < b)
{
System.Console.Write("Condition is true");
}
return b;
}

Here is the MSIL code generated by Visual Studio 2005 compiler.

.method public hidebysig instance int32 IfStructure(int32 a, int32 b) cil managed
{
// Code size 31 (0x1f)
.maxstack 2
.locals init ([0] int32 CS$1$0000,
[1] bool CS$4$0001)
IL_0000: nop
IL_0001: ldarg.1
IL_0002: ldarg.2
IL_0003: clt
IL_0005: ldc.i4.0
IL_0006: ceq
IL_0008: stloc.1
IL_0009: ldloc.1
IL_000a: brtrue.s IL_0019

IL_000c: nop
IL_000d: ldstr "Condition is true"
IL_0012: call void [mscorlib]System.Console::Write(string)
IL_0017: nop
IL_0018: nop

IL_0019: ldarg.2
IL_001a: stloc.0
IL_001b: br.s IL_001d
IL_001d: ldloc.0
IL_001e: ret
} // end of method ControlStructures::IfStructure

We now start parsing:
-------------------------------------------------------------------------------------
.method public hidebysig instance int32 IfStructure(int32 a, int32 b) cil managed
{
// Code size 31 (0x1f)
.maxstack 2
.locals init ([0] int32 CS$1$0000,
[1] bool CS$4$0001)
--

These lines generate output that we do without any processing of MSIL. There is method definition and local variables. We change the local variable names to C# current names without conflict. For simplicity here we just replace '$' with '_'. One thing to evaluate the MSIL instructions we must keep a map of local variables with variable number. In this example CS$1$0000 is local variable 0 of type int. For clarity we do not show the map here. A simple STL map should work.

Output:
public int IfStructure(int a, int b)
{
int32 CS_1_0000;
bool CS_4_0001;

Stack: [Empty]

-------------------------------------------------------------------------------------
IL_0000: nop
IL_0001: ldarg.1
IL_0002: ldarg.2
--
So we push method argument 1 and 2 to stack:

Output: [None]

Stack:
a,b

-------------------------------------------------------------------------------------
IL_0003: clt
--

This instructs us if stack top-1 is less than stack top. The two elements are popped from stack and result goes to stack. No output of course.

Output: [None]

Stack:
a<b

-------------------------------------------------------------------------------------
IL_0005: ldc.i4.0
--

Load (aka push) constant integer of value 0 on stack.

Output: [None]
Stack: a<b,0
-------------------------------------------------------------------------------------
IL_0006: ceq
--

Check if stack top-1 equals stack top. Result goes to stack.
Output: [None]
Stack: a < b == 0

-------------------------------------------------------------------------------------
IL_0008: stloc.1
IL_0009: ldloc.1
--

What else? Store stack top in local variable 1 and load that on tack again. We decided previously when we store some value in a local variable we use assignment to that variable and output that code. For clarity I added parentheses:

Output:CS_4_0001 = (a < b == 0)
Stack:CS_4_0001

-------------------------------------------------------------------------------------
IL_000a: brtrue.s IL_0019
--

We have got a conditional branch. We create a block starting from here to IL_0019. And put them in curly braces. And our condition is on the stack. We find a true condition branch so we negate it and put it as if structure as I said at the beginning.

Output:
if(!CS_4_0001)
{
IL_000c: nop
IL_000d: ldstr "Condition is true"
IL_0012: call void [mscorlib]System.Console::Write(string)
IL_0017: nop
IL_0018: nop
}

Stack: [Empty]

-------------------------------------------------------------------------------------
IL_0019: ldarg.2
IL_001a: stloc.0
IL_001b: br.s IL_001d
IL_001d: ldloc.0
IL_001e: ret

We do not parse them here. They are very simple to understand.
===========================================================

Ok we now can work on little more complex code3s than easiest. This will also produce codes that was generated by "for structure" but in a funny way. If we add a little more intelligence to produce "goto" output code for special branching that we cannot handle with if, we get following result.

The code like this:

for(int i=0;i<10;i++)
{
...
}
...

will be converted to:

int i;
i=0;
label_1:

if(i<10)
{
---
i++;
goto label_1;
}
...

For is not our today’s material. We'll look at loops next time.

=====================================================================================
Now you may find that our theory generates a funny code block like:

CS$4$0001=((a<b)==0);

if(!CS$4$0001)
{

---

}

Here the optimization comes to scene. But we skip them for future.