Showing posts with label computer architecture. Show all posts
Showing posts with label computer architecture. Show all posts
Tuesday, August 17, 2010
Key Computer Architecture Techniques - Pipelining
Pipelining
It is an implementation technique where multiple instructions are overlapped in execution. The computer pipeline is divided in stages. Each stage completes a part of an instruction in parallel. The stages are connected one to the next to form a pipe - instructions enter at one end, progress through the stages, and exit at the other end.
Because the pipe stages are hooked together, all the stages must be ready to proceed at the same time. We call the time required to move an instruction one step further in the pipeline a machine cycle . The length of the machine cycle is determined by the time required for the slowest pipe stage.
For example, the classic RISC pipeline is broken into five stages with a set of flip flops between each stage.
1. Instruction fetch
2. Instruction decode and register fetch
3. Execute
4. Memory access
5. Register write back
Pipelining does not help in all cases. There are several possible disadvantages. An instruction pipeline is said to be fully pipelined if it can accept a new instruction every clock cycle. A pipeline that is not fully pipelined has wait cycles that delay the progress of the pipeline.
Advantages of Pipelining:
1. The cycle time of the processor is reduced, thus increasing instruction issue-rate in most cases.
2. Some combinational circuits such as adders or multipliers can be made faster by adding more circuitry. If pipelining is used instead, it can save circuitry vs. a more complex combinational circuit.
Disadvantages of Pipelining:
1. A non-pipelined processor executes only a single instruction at a time. This prevents branch delays (in effect, every branch is delayed) and problems with serial instructions being executed concurrently. Consequently the design is simpler and cheaper to manufacture.
2. The instruction latency in a non-pipelined processor is slightly lower than in a pipelined equivalent. This is due to the fact that extra flip flops (pipeline registers/buffers) must be added to the data path of a pipelined processor.
3. A non-pipelined processor will have a stable instruction bandwidth. The performance of a pipelined processor is much harder to predict and may vary more widely between different programs.
Labels:
computer architecture,
computer science,
jobs,
pipeline,
pipelining
Friday, August 13, 2010
Computer Performance
From Stanford EE282 + Computer Architecture: A Quantitative Approach, 4th Edition
CPUTime = Seconds/Program
= Cycles/Program * Seconds/Cycle
= Instructions/Program * Cycles/Instruction * Seconds/Cycle
= IC * CPI * CCT
IC: Instruction Count
CPI: Clock Cycles Per Instruction
CCT: Clock Cycle Time
Amdahl’s Law
Speedup = Execution time for entire task without using the enhancement/Execution time for entire task using the enhancement when possible
It should be greater than 1 (when there is an improvement, that is)
new_execution_time
= (original_execution_time * (1 - enhanced_fraction)) + original_execution_time * enhanced_fraction * (1 / speedup_enhanced)
= (original_execution_time)((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
speedup_overall
= original_execution_time / new_execution_time
= 1/((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
In the case of parallelization, Amdahl's law states that if enhanced_fraction is the proportion of a program that can be made parallel (i.e. benefit from parallelization), and (1 − enhanced_fraction) is the proportion that cannot be parallelized (remains serial), then the maximum speedup that can be achieved by using N processors (N times faster for the part that can be enhanced = speedup_enhanced) is
speedup_overall
= 1/((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
In the limit, as N tends to infinity, the maximum speedup tends to 1 / (1 − enhanced_fraction).
- Latency or execution time or response time
- Wall-clock time to complete a task
- Bandwidth or throughput or execute rate
- Number of tasks completed per unit of time
CPUTime = Seconds/Program
= Cycles/Program * Seconds/Cycle
= Instructions/Program * Cycles/Instruction * Seconds/Cycle
= IC * CPI * CCT
IC: Instruction Count
CPI: Clock Cycles Per Instruction
CCT: Clock Cycle Time
Amdahl’s Law
Speedup = Execution time for entire task without using the enhancement/Execution time for entire task using the enhancement when possible
It should be greater than 1 (when there is an improvement, that is)
new_execution_time
= (original_execution_time * (1 - enhanced_fraction)) + original_execution_time * enhanced_fraction * (1 / speedup_enhanced)
= (original_execution_time)((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
speedup_overall
= original_execution_time / new_execution_time
= 1/((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
In the case of parallelization, Amdahl's law states that if enhanced_fraction is the proportion of a program that can be made parallel (i.e. benefit from parallelization), and (1 − enhanced_fraction) is the proportion that cannot be parallelized (remains serial), then the maximum speedup that can be achieved by using N processors (N times faster for the part that can be enhanced = speedup_enhanced) is
speedup_overall
= 1/((1-enhanced_fraction) + enhanced_fraction/speedup_enhanced)
In the limit, as N tends to infinity, the maximum speedup tends to 1 / (1 − enhanced_fraction).
Friday, August 6, 2010
System Diagram of A Modern Laptop
From Wikipedia Intel X58
- Intel X58: Intel X58 Chipset
- QPI: The Intel QuickPath Interconnect is a point-to-point processor interconnect developed by Intel to compete with HyperTransport
- I/O Controller Hub (ICH), also known as Intel 82801, is an Intel southbridge on motherboards with Intel chipsets (Intel Hub Architecture). As with any other southbridge, the ICH is used to connect and control peripheral devices.
- EHCI: The Enhanced Host Controller Interface (EHCI) specification describes the register-level interface for a Host Controller for the Universal Serial Bus (USB) Revision 2.0.
- DMI: Direct Media Interface (DMI) is point-to-point interconnection between an Intel northbridge and an Intel southbridge on a computer motherboard. It is the successor of the Hub Interface used in previous chipsets. It provides for a 10Gb/s bidirectional data rate.
- LPC: The Low Pin Count bus, or LPC bus, is used on IBM-compatible personal computers to connect low-bandwidth devices to the CPU, such as the boot ROM and the "legacy" I/O devices (behind a super I/O chip).
- SPI: The Serial Peripheral Interface Bus or SPI (pronounced "ess-pee-i" or "spy") bus is a synchronous serial data link standard named by Motorola that operates in full duplex mode. Devices communicate in master/slave mode where the master device initiates the data frame.
- Intel Matrix Storage Technology: It provides new levels of protection, performance, and expandability for desktop and mobile platforms. Whether using one or multiple hard drives, users can take advantage of enhanced performance and lower power consumption. When using more than one drive, the user can have additional protection against data loss in the event of a hard drive failure.
- Intel Turbo Memory with User Pinning: An on-motherboard flash card, Intel's Turbo Memory is designed to act as another layer in the memory hierarchy, caching data where possible and improving performance/battery life in notebooks. User pinning offers more options to the user to improve system applications launch time and responsiveness.
Labels:
computer architecture,
diagram,
Intel,
Laptop,
Processor
Subscribe to:
Posts (Atom)
