Friday, September 3, 2010

SCSI Interconnect

The original SCSI interconnect was implemented with 8 data lines and a few control lines. The parallel bus architecture evolved to higher bandwidth via wider data path and higher clock speed. The signal skew problem limited the distance of the parallel SCSI interconnect. Also, each daisy chained SCSI string is limited to 16 SCSI ID, which caps the capacity.

For direct attached SCSI configuration, SAS addresses the limitation of parallel SCSI.

SCSI Architecture

SAM-2 (SCSI Architecture Model) defines the relationship between initiator and target. Serial SCSI implementation such as Fibre Channel, serial-attached SCSI (SAS) and iSCSI are a component of the SAM-2 definition for SCSI-3 commands.

The client-server requests and responmses are exchanged across some form of physical transport governed by SCSI-3 service deliveryt protocol such as FC or iSCSI.

Read/Write of data are performed with a series of SCSI commands, delivery requests, delivery actions and responses. SCSI commands and parameters are specified in the Command Descriptor Block (CDB). CDB is encapsulated in the FCP IU (information unit).

The operating system views a WRTIE operation as a single operation but underneath, there are multiple SCSI exchanges:

(1) WRITE trigger the creation of a client in the initiator.
(2) Initiator in turn issues a SCSI command request to the target to prepare a buffer.
(3) The target device server issues a delivery action request when the buffer is ready.
(4) The initiator sends data block
(5) After the data is recieved, the target sends another delivery action request to ask for anotherdata block.
(6) When all data blocks have been recieved, the target sends a Response to mark the end of the WRITE operation

For READ operation, the request and response directions are reversed, with the host prepare buffer for data from the disk.

SCSI LUN

SCSI (Small Computer System Interface) is implemented in a client/server model. Computer acts as client (initiator) and storage device acts as server (target). The SCSI command processing entity within the storage target represents a logical unit and assigned a number (LUN). SCSI targets are assigned a 3-part bus/target/LUN descriptor. Bus refer to one of the several SCSI that installed on the hot (e.g. HBA, iSCSI network card). Bus supported multiple daisy-chained disks (target). LUN represents a SCSI server within the target.

Operating systems expect to boot from LUN 0. For multiple server to boot from the same disk array, the array needs to support multiple LUN 0s. This is done by mapping the actual LUN to virtual LUN. LUN mapping also include LUN masking used to isolate set of LUN from other hosts or users.

Monday, December 7, 2009

Interrupt

Maskable interrupt can be generated by hardware or software by asserting the INTR line. They are maskable because programmer can disable the processor from recgonizing the INTR signal or disable the interrupt controller from accepting the interrupt request from selected device.

Non-Maskable interrupt (NMI) is generated by the chipset when serious hardware problem was detected in the system board. The processor's NMI input is asserted.

Software exception

Software exception refers to the problem when executing an instruction or its operands. The processor attempts to recovery gracefully by invoking a special exception handler.

A fault is an exception reported at the start of the instruction that caused the exception. The instruction can be restarted after the handler fixes the problem (e.g. page fault). A trap is an exception reported after the offending instruction has been executed. An abort does not always reliably supply the instruction that caused the problem. This makes it impossible for the exception handler to fix the problem and resume program execution.

Demand Paging in 386

Segmentation complicates programming. Paging can be used to present a flat 32-bit (4GB) address space and yet provide protection among tasks. There is no way to switch off segmentation in the processor. However, if all segments defined in GDT was set to R/W, start at 00000000h and 4GB in legnth, segmentation is effectively eliminated.

Paging is enabled by setting PG bit in CR1 to 1. The paging unit intercepts all 32-bit linear memory addresses generated by the segment unit and perform a redirection mapping using a 2-level page table structure.

The top 10 bits of the linear address is used to index into the page directory to yield the base address of the Page table. The next 10 bits of the linear address is then used to index into the Page Table to yield the start address of the memory page. The last 12 bits is then used to access the location as offset.

If the Page Table was not in memory, a page fault is triggered. The linear address is latched into CR2 so that it could be accessed by the OS's page fault handler. When the target page is not in memory, similar action is performed.

TLB (Translation Lookaside Buffer) is used to short-circuit the look up. The segment unit sends the linear address to both TLB and Paging Unit. The top 20-bits of the linear address is compared against the cached entries in the TLB. If a match is found, it will disable the paging unit and send the corresponding 20-bits physical address mapping onto the FSB for retrieval.

Segmentation in IA32 Real Mode

The start address of the segment must be within the first 1M memory space. The length of segment is fixed to 64KB (most significant 16 bits in the address). There is also no protection of segment among tasks.

The 64K memory above 1M line is called extended memory or HMA (High memory area). In real mode, one can access HMA by setting the segment register to xFFFF. For example,

mov ax,ffff
mov ds,ax
mov al,[0010]

The effective address - FFFF0+0010=10000.

This method is not effective for 8086/8088 which does not have A20 pin. Instead, the address will be wrapped around to 0000 (segment wrapping).

Monday, November 30, 2009

Memory Alignment for 386

When the processor initiates a transaction on the FSB, the logic external to the processor takes the least 2 significant bits of the address as always zero. In other word, the processor can only address memory locations at Dword boundary. The processor implement 4 output pins (BE0# to BE3$) instead to address individual byte in the Dword. (Each BE pin is used to select a separate memroy bank DIMM?) For example, for location zero of the Dwrod, the processor asserts BE0# pin and the target byte will be output over data path 0 (D[7:0]). For location three of the Dword, the processor asserts BE3# pin and the data will be output over data path 3 (D[31:24]).

To execute this instruction - mov eax,[0101], processor will need to access address locations at 0100 and 0104 (Dword boundary). Then it extracts the last 3 bytes from 0100 and the first byte from 1004 to form the final Dword to be loaded in eax. This degrades performance. Moreover, it could further trigger double cache misses or page misses. Therefore, Dword alignment of data is important. Starting from 486, all IA32 procssors will flag out this condition (Aligment Check Exception).