Monday, September 21, 2009

Useless mov edi,edi in the Prologue

The seemingly useless statement is used to enable hot patching (patching without stopping the component). The 2-byte instruction can be changed to a short jmp operation (within a range of 127 bytes in either direction). To extend the jmp target, NOP statements are generated before the function labels so that a long jmp statement could be patched in:

xor eax,eax
jmp xyz
nop
nop
nop
nop
nop
func-abc:
mov edi,edi
push ebp
mov ebp,esp
:

Frame Point Omission

FPO is an optimization technique. ebp is used as a general purpose register rather than the stack frame base pointer. Execution is sped up by the availability of this additional register.

Call Convention

Stdcall pushes the argument from right to left onto the stack. The called function is responsible to remove the parameters passed in by decrementing the esp by the length of the parameters. 

Cdecl call convention differs from Stdcall by having the calling function to remove the argument passed from the stack. 

Stdcall is preferred because the clean up is done one place (no mater how many times it is being called), which is simpler. Cdecl is used for C/C++ because they support variable number of parameters for function call. As the called function will not know the number of parameters beforehand, the clean up has to be performed by the calling function instead.

Linker generates special name for different call conventions. For Stdcall, function name will be prefixed by "_" and appended by "@", follow by the number of bytes of stack space required. For Cdecl, function name is prefixed by "_".

Fastcall uses ecx and edx to pass the first 2 argument. Clean up is by the called function, similar to Stdcall. Function name is prefixed by "@", appended by "@" and followed by the number of bytes of stack space required.

Thiscall passes this point via exc and the rest of arguments on the stack. Clean up is by called function.

Stack Frame

Before calling a function, the caller will first reserved space for the parameters in the stack. For example, assuming the parameters occupies 20 bytes:

sub esp,14h

Following this is a series of mov statment to move the parameters to the stack using offset with ebp. For example,

mov dword ptr [edp-14h],3
:
:

Then the call operation is used to jump to the function. Call will push the eip onto the stack (esp will advance as a result).

At the beginning of the function, the compiler generates a stack frame using the frame base pointer register ebp. The function prologue saves the current ebp onto the stack before setting up a new stack frame:

mov edi, edi
push ebp
mov edp, esp

As a result, the ebp of the new frame points to the old ebp value (the last frame base). The ebp is then used to access the parameter (positive offset) and local variables (negative offset).

At this stage, the call stack contains the following (growing downwards):

parm1
parm2
:
return address
saved ebp of caller
local variable1
local variable2
:

When the function finishes, the epilogue restore the previous stack frame

add esp,14h ; clean up the parameter stack space assuming this is stdcall
mov esp, ebp
pop ebp
ret

Saturday, August 22, 2009

SLIP and PPP

SLIP (Serial line IP) is a link level protocol to carry IP over serial line (e.g. modem, RC232 interace). Each IP datagram is terminated by END character. The END character will be ESCAPEd if presents in the datagram. There is no checksum in SLIP and error detection and handling is assumed to be done in the upper protocol layer.

As serial interface speed is typically slow, CSLIP is used to optimize throughput by reducing the 20+20 TCP+IP header to 3 or 5 bytes. This is done by maintaining state information (fields that rarely changed) in the CSLIP protocol thus removing the need for such information to be present in the normal TCPIP header. CSLIP can maintain up to 16 connections. The smaller header size improve response time for interactive session on serial line.

PPP (Point-to-point Protocol) improves on SLIP by adding LCP (Link Control Protocol) and NCP (Network Control Protocol) capability. LCP allow the both ends to negotiate options (e.g. IP address negitiation for both ends). NCP allows PPP to support more than 1 network protocol (i.e. no just IP) on one serial line. Finally, PPP included checksum.

Sunday, August 9, 2009

Challenges to Pipleining and Superscalar

Data Hazard refers to the use of related data in 2 instruction that prevent them from executing simultaneously. For example, the output of the one instruction is used as an input to the next instruction. Pipelined processors use "forwarding" to resolve this issue. Output port of the ALU is fed into the input port directly and bypassing the register-file write stage. Superscalar processor uses "register renaming" to decouple instructions using the same register in the calculation. For example, the following 2 instruction can be executed simultanously using register renaming technique.

Add A, B, C; add a and b and store result in c
Add D, B, A; add d and b and store result in a

Structure Hazard refer to the shortage of resources to execute multiple instruction simultaneously. In a superscalar design, it takes a large number of wire to connect each ALU to the register. Hence, CPU registers are grouped into a special unit called register file. Register files are like memory array which consists a data bus and 2 ports - read and write ports. for example, ALU accesses the register file's read port and requests the data to be placed on the bus. A single read port allows the ALU to access a signle registr at a time. Therefore, for 3 operand instruction like the above requires 2 read port and 1 write port. Modern CPU also uses separate regiester files to store integer, floating-point and vector numbers as each of them uses separate execution units. Another reason for this separation is to keep the register file size small. The large the register file, the slower the access will be.

Control (Branch) Hazard arises when the processor arrives at a conditional branch instruction. Branch prediction is used to get around this type of stall. Instruction cache is used to improve the performance for loading the next instruction from a branch.

ISA

In 1960, IBM S/360 introduced the concept of ISA as a layer of abstraction to the underlining CPU hardware microarchitecture. Programs written on an ISA are guaranteed to run on any CPU that implement the ISA. ISA provides a standardized way to expose the features of a system's hardware that allows manufactures to enhance the implementation without breaking programs. ISA is implemented using microcode engine, wich consists of some storage, microcode ROM which holds the microcode programs, and an execution unit that translate the standard instruction to the ones specific to the hardware implementaiton.

The drawback of microcode engine is it is slower than direct decoding. (Modern microcode engine has approached 99% of the speed.) However, the benefit of abstraction is so signifcant that outweight this slight penalty.