Linux Kernel Interrupt Handling: From Hardware Exceptions to Softirq, a Deep Dive into the Complete Mechanism
Table of Contents
- Overview of the Interrupt Subsystem
- Hardware Foundation: Interrupt Controllers and IRQ Lines
- Interrupt Descriptor Table (IDT) and IDT Initialization
- Complete Path of Hardware Interrupt Handling
- Top-Half and Bottom-Half: The Origin of Interrupt Splitting
- Softirq: The Kernel's Low-Priority Interrupt Mechanism
- Tasklet: A Bottom-Half Mechanism Built on Softirq
- Workqueue: Bottom-Half Execution in Process Context
- Threaded IRQ: The Modern Approach to Interrupt Handling
- Timer Interrupts and Kernel Time Management
- Performance Monitoring and Tuning Tools
- Summary
1. Overview of the Interrupt Subsystem
Interrupts are the core mechanism by which the CPU responds to hardware events in modern operating systems. Whether it is a keyboard keystroke, a network packet arrival, or a disk I/O completion, all notifications to the CPU are delivered through interrupts. In the Linux kernel, the interrupt subsystem is a complex and critical component that spans hardware abstraction, context switching, scheduling, and concurrency control.
The entire interrupt handling flow can be divided into several key stages:
- Hardware Interrupt Trigger: A peripheral device sends an interrupt signal to the interrupt controller via an IRQ line.
- Interrupt Controller Routing: The interrupt controller (APIC/IO-APIC) routes the interrupt to a specific CPU core based on priority and affinity.
- CPU Response: The CPU saves the current execution context and jumps to the corresponding interrupt handler.
- Top-Half Handler Execution: Time-critical operations are executed in interrupt context, such as acknowledging the interrupt and reading hardware status.
- Bottom-Half Handler Execution: Non-time-critical, time-consuming operations are deferred for later execution.
- Context Restore: Restore the previously saved context and return to the pre-interrupt execution state.
2. Hardware Foundation: Interrupt Controllers and IRQ Lines
2.1 8259A PIC (Progrepesmable Interrupt Controller)
The early x86 architecture used the Intel 8259A PIC chip to manage hardware interrupts. A single 8259A can manage 8 interrupt lines (IRQ0-IRQ7), and through a master-slave cascading approach, it can be extended to 15 interrupt lines (IRQ0-IRQ15). The 8259A uses edge-triggered or level-triggered interrupt signals, with fixed priority (IRQ0 has the highest priority).
2.2 APIC (Advanced Programmable Interrupt Controller)
As the number of CPU cores increased, the traditional 8259A architecture encountered severe bottlenecks in multi-processor environments. The APIC architecture arose at the historic moment and mainly consists of two parts:
- Local APIC (LAPIC): Integrated into each CPU core, responsible for receiving and processing interrupts directed to the local CPU. LAPIC includes a timer, performance counters, and interrupt registers.
- I/O APIC: Usually integrated into the South Bridge chip, responsible for receiving interrupt signals from peripheral devices and distributing them to the appropriate LAPIC message interrupts. Modern I/O APICs typically support 24 interrupt lines (IRQ0-IRQ23).
2.3 MSI and MSI-X
MSI (Message Signaled Interrupts) is a modern interrupt mechanism introduced in PCI 2.2. Unlike traditional pin-based interrupts, MSI issues interrupt requests by writing specific data to a specific memory address hosted by the LAPIC. This approach avoids the sharing and competitive interrupt delay problems of physical IRQ lines and supports the simultaneous allocation of multiple interrupt vectors (MSI-X supports up to 2048 interrupt vectors).
2.4 Interrupt Affinity and Interrupt Balancing
In multi-core systems, interrupt affinity (Interrupt Affinity) determines which CPU cores a specific interrupt can be processed by. By setting /proc/irq/IRQ#/smp_affinity, users can bind specific interrupts to specific CPU cores to optimize cache locality and load balancing. irqbalance is an interrupt auto-balancing daemon that distributes interrupts to different CPU cores to avoid the lethality of a core becoming overloaded.
3. Interrupt Descriptor Table (IDT) and IDT Initialization
3.1 IDT Structure
The Interrupt Descriptor Table (IDT) is the basis for the x86 architecture to implement interrupt handling. It is essentially an array of 256 entries (on x86_64), with each entry being an 8-byte (32-bit mode) or 16-byte (64-bit mode) descriptor containing the following information:
- Handler function address (offset)
- Target code segment selector (segment selector)
- Gate type (Interrupt Gate / Trap Gate / Task Gate)
- Privilege level (DPL)
- Present bit (P)
The differences between Interrupt Gate and Trap Gate lie in interrupt masking behavior: Interrupt Gate will clear the IF flag to disable interrupts when entering the handler, preventing nested interrupts; Trap Gate does not modify the IF flag, allowing interrupts to be nested.
3.2 IDT Initialization Process
In the Linux kernel, IDT initialization is divided into two stages:
Stage 1: Early IDT Setup (setup_arch() → idt_setup_early_traps())
Initialized in the start_kernel() phase, this stage establishes a minimal IDT to handle critical exceptions (such as #DF Double Fault, #GP General Protection Exception, #PF Page Fault). At this point, memory management and scheduling systems are not yet running, and must use a statically allocated IDT.
Stage 2: Complete IDT Setup (idt_setup_traps())
Initialized after kernel startup is complete. It will:
- Register a unified interrupt entry point for 256 vector numbers through the
SETGATE()macro. - Register interrupt handling stubs for system calls (vector 0x80 on 32-bit, MSR_LSTAR on 64-bit
syscall). - Register entries for local APIC interrupts (such as SPURIOUS_APIC_VECTOR, timer interrupt, etc.).
3.3 Interrupt Vector Number Allocation
On the x86_64 architecture, the Linux kernel's interrupt vector number allocation is as follows:
| Vector Range | Use |
|---|---|
| 0-31 | CPU exceptions (Non-maskable faults, traps, aborts) |
| 32-127 | External hardware interrupts (via IOAPIC/MSI) |
| 128 | System call entry point (legacy syscall) |
| 129-238 | Dynamic allocation (MSI-X, IPIs) |
| 239-255 | Special vectors (LOCAL_TIMER_VECTOR, SPURIOUS_APIC_VECTOR, etc.) |
4. Complete Path of Hardware Interrupt Handling
4.1 Hardware Automatically Performed Actions
When the CPU receives an interrupt, it automatically executes the following operations (in hardware order):
- Push the RFLAGS, CS, and RIP registers onto the current stack.
- If the interrupt occurs at a different privilege level (e.g., from user mode to kernel mode), also push SS and ESP.
- Clear the IF flag (if the interrupt uses an Interrupt Gate) to disable external maskable interrupts.
- The target address is found based on the interrupt vector number and IDT entry.
- Jump to the entry point of the interrupt handler.
4.2 Entry Point for the Interrupt Handling Stub
On x86_64, each interrupt vector corresponds to an assembly entry point registered by the interrupt macro in arch/x86/entry/entry_64.S. These entry points perform the following tasks:
- Save the general-purpose registers (RAX, RCX, RDX, etc.) to the kernel stack.
- For interrupts with error codes (e.g., #DF, #PF), process the error code.
- Call the high-level C interrupt handler through the
idtentrymacro.
4.3 Call Flow of High-Level Interrupt Handling
The flow after entering high-level interrupt handling is as follows:
do_IRQ()
→ handle_irq_event()
→ handle_irq_event_percpu()
→ __handle_irq_event_percpu()
→ handler->action->handler() // Call a series of interrupt action chains registered by the device driver
In do_IRQ(), the kernel first calls the processing function of the corresponding interrupt descriptor's irqchip (such as mask_ack_irq()) to mask and confirm the interrupt source, then traverses the handler list registered on that IRQ line, and calls each handler in sequence.
4.4 irqdesc and irqaction Data Structures
The interrupt descriptor struct irq_desc is the central data structure for the kernel to manage interrupts, containing:
- irq_common_data: Affinity, CPU mask, and parent hardware interrupt controller pointer.
- irq_data: Hardware interrupt number, domain (irq_domain), and hardware controller information.
- kstat_irqs: Per-CPU interrupt statistics.
- action: Pointer to the interrupt handler list (
struct irqaction). - depth: Interrupt disable nesting count.
- lock: Spinlock to protect concurrent access.
struct irqaction represents a registered interrupt handler which contains:
- handler: The interrupt handler function pointer (i.e.,
irq_handler_t). - dev_id: Device identifier for shared interrupt lines.
- irq: Interrupt number.
- flags: Configuration flags such as IRQF_SHARED, IRQF_TRIGGER_RISING.
5. Top-Half and Bottom-Half: The Original Intent of Interrupt Splitting
5.1 Why Interrupt Splitting Was Needed
The original interrupt handling in the Linux kernel executed all processing directly in interrupt context. However, in practice, hardware interrupt handlers often require a large amount of data processing, such as network packet encapsulation parsing or filesystem metadata updates. Executing these time-consuming tasks in interrupt context leads to two serious problems:
- Interrupt Loss: When an interrupt is being handled, the same interrupt is typically masked (masked). If the handler spends too long, the device issues a new interrupt (such as a network card receiving a new packet) during this period will be lost.
- System Response Delay: Interrupt context disables the kernel's preemption capability (in some cases even interrupt nesting). If a large number of time-consuming operations are executed in interrupt context, the system's response time will significantly increase.
5.2 The Difference Between Top-Half and Bottom-Half
To solve this problem, Linux 2.3 introduced the interrupt handling splitting mechanism: interrupt handlers are split into a "Top-Half" (top half) that runs in interrupt context and a "Bottom-Half" (bottom half) that runs after the interrupt exits.
- Top-Half: The handler directly registered by the device driver via
request_irq(). It executes in interrupt context (interrupt context/disabled interrupt), performing only time-critical tasks (such as reading hardware status registers, acknowledging interrupts, copying data to a kernel buffer). Execution time must be extremely short. - Bottom-Half: Deferred execution mechanism, executes in a softer interrupt context or process-like context (depending on the specific mechanism), processing time-consuming tasks (such as protocol stack processing, file I/O, etc.). While the bottom half executes, it allows other interrupts to be responded to normally.
6. Softirq: A Kernel Low-Priority Interrupt Mechanism
6.1 Design Concept of Softirq
Softirq is one of the earliest preemptive (non-trappable) kernel mechanisms in the Linux kernel. Its core concept is: simulating a software-implemented interrupt to execute interrupt bottom-half tasks outside of hardware interrupt context. Softirq runs in a "soft interrupt" context, which allows hardware interrupts to be re-enabled while prohibiting the execution of higher-priority softirqs.
6.2 Priority and Types of Softirq
The Linux kernel defines 10 softirq types (since 5.10, the actual active types are less than 10), with the priority determined by the index in the enumeration enum list:
enum {
HI_SOFTIRQ=0, // High-priority softirq (tasklet_hi)
TIMER_SOFTIRQ, // Timer softirq
NET_TX_SOFTIRQ, // Network transmission
NET_RX_SOFTIRQ, // Network reception -- the most commonly used
BLOCK_SOFTIRQ, // Block device
IRQ_POLL_SOFTIRQ, // Interrupt polling
TASKLET_SOFTIRQ, // Tasklet (common)
SCHED_SOFTIRQ, // Scheduler
HRTIMER_SOFTIRQ, // High-resolution timer
RCU_SOFTIRQ, // RCU garbage collection -- usually the highest priority
};
6.3 Registration and Triggering of Softirq
Softirq usage follows a statically registered model:
- Register a handler functions for the specified
softirq_vecviaopen_softirq()(can only be registered at compile time or during module load, not dynamically). - Trigger the pending flag of the specified softirq by raising a
__raise_softirq_irqoff(). Location is a CPU-specific unsigned integerirq_stat.__softirq_pending.
6.4 Execution Timing of Softirq
Softirq does not have an independent scheduling unit but is instead executed through the following trigger points:
- Hardware interrupt return path:
irq_exit()indo_IRQ(). If there are pending softirqs, invokeinvoke_softirq(). - ksoftirqd kernel thread: When the softirq trigger frequency is too high (if not processed within a certain time), the ksoftirqd/ kernel thread is awakened to execute softirq in SCHED context. This is to prevent starvation of the user processes caused by the time-consuming full occupation softirq.
- Explicit check points: Places where a context switch may occur in the kernel (such as
cond_resched()when the kernel is preemptible).
6.5 Softirq Limitations
- Static Registration: Cannot be dynamically registered/removed, and must be defined through macros at compile time to ensure initialization.
- Strict Synchronization Requirements: Softirq handlers of the same type can execute synchronously on different CPUs, which means shared data must be protected strictly by per-CPU variables or lock mechanisms.
- Exclusive Execution: The same softirq type cannot execute simultaneously on different CPUs (by default, only when the handler explicitly invokes a rest of function triggers ksoftirqd can it be re-entrant).
7. Tasklet: A Bottom-Half Mechanism Built on Softirq
7.1 Design Goals of Tasklet
Tasklet is a higher-level mechanism built on top of softirq, designed to provide a simpler interface for device drivers to achieve bottom-half handling. Tasklet is internally implemented using two softirq numbers: TASKLET_SOFTIRQ (normal priority) and HI_SOFTIRQ (high priority).
7.2 Tasklet Data Structure
struct tasklet_struct {
struct tasklet_struct *next; // Pointer forming a linked list
unsigned long state; // TASKLET_STATE_SCHED / TASKLET_STATE_RUN
atomic_t count; // Disable counter
void (*func)(unsigned long); // Handler function
unsigned long data; // Parameters passed to the handler
};
Each CPU maintains two tasklet linked lists: tasklet_vec (normal) and tasklet_hi_vec (high priority). When tasklet_schedule() is called, the tasklet is inserted into the corresponding linked list and the softirq of TASKLET_SOFTIRQ is raised.
7.3 Tasklet Scheduling Rules
- A scheduled tasklet will execute only once on the same CPU.
- If the tasklet executes on one CPU, its
TASKLET_STATE_SCHEDlocation is signaled, even iftasklet_schedule()is called again before execution, it is guaranteed to be executed only once. - Same tasklets cannot run simultaneously on different CPUs (compared to softirq, this is a stricter exclusivity convention).
7.4 Differences Between Tasklet and Softirq
| Features | Softirq | Tasklet |
|---|---|---|
| Registration Method | Static (compile time) | Static or Dynamic (dynamic via tasklet_init) |
| Concurrent Execution | Same type can execute synchronously across multiple CPUs | Same types cannot execute synchronously across multiple CPUs |
| Use Scenarios | High-frequency, high-throughput (network, block) | General-purpose device drivers |
| Overhead | Lower | Slightly higher (linked list operations) |
8. Workqueue: Bottom-Half Execution in Process Context
8.1 Design Goals of Workqueue
Unlike softirq and tasklet executing in interrupt context, workqueue allows bottom-half tasks to execute in a specialized kernel thread (worker thread) in user-like (process context). The most significant advantage of this is: workqueue context allows blocking operations (mutex acquisition, memory I/O operations, sleep, etc.), so workqueue has much higher flexibility and complexity in handling long-running tasks.
8.2 Workqueue Architecture
The workqueue system is primarily organized into three layers:
- API Layer: Provides
schedule_work(),schedule_delayed_work(), and other interfaces. - cmwq (Concurrency Management Workqueue): Manages a pool of worker threads on each CPU and automatically adjusts the number of worker threads based on the system's load (Dynamic worker pool since Linux 5.10 may be replaced by a new bound worker pool).
- Worker Thread (worker thread):
kworker/n:kkernel threads in user-like context, executing actual work function functions from the pending queue.
8.3 Worker Thread Management Mechanism
| State | Description |
|---|---|
| IDLE | Worker thread is blocked and not executing any work. A maximum of 2 IDLE workers are retained on each CPU by default unless conversion is busy. |
| BUSY | Worker thread is executing a work function. |
| SYSTEM | When there is a long-blocking work in the queue, a new worker thread is dynamically created to handle it to avoid starving the current worker pool. |
8.4 How to Choose Between Workqueue and Tasklet/Softirq?
- If the bottom-half task needs to block/sleep: use workqueue.
- If the bottom-half task needs extremely high throughput and frequency (e.g., network packet capture) and does not block: use softirq or tasklet.
- If the bottom-half task has no special requirements: prioritize tasklet.
9. Threaded IRQ: A Modern Way of Interrupt Handling
9.1 Introduction to Threaded IRQ
Since Linux 3.1, kernel developers introduced request_threaded_irq() API, which allows device drivers to split interrupt handlers into two parts: hardirq handler (hard interrupt handler) and threaded handler (thread handler). The hardirq handler plays the role of top-half, executing in interrupt context, and quickly checking whether the interrupt originates from the device being driven; it returns IRQ_WAKE_THREAD, then the threaded handler executes in a dedicated kernel thread to complete complex business logic.
// Request a threaded interrupt
request_threaded_irq(irq, hardirq_handler, threaded_handler, flags, name, dev);
// hardirq_handler optionally returns IRQ_WAKE_THREAD to wake up the thread
9.2 Handling Flow of Threaded IRQ
- An interrupt occurs, executes the hardirq handler (interrupts disabled).
- If the hardirq handler returns
IRQ_NONE, it indicates the interrupt is not from this device and the next handler continues. - If it returns
IRQ_WAKE_THREAD, it wakes up the corresponding thread handler. - The thread executes
thread_fnin SCHED context, which may block and sleep. - If the hardirq handler returns
IRQ_HANDLED, processing is complete and the thread is not woken.
9.3 Advantages of Threaded IRQ
- No need to manually operate tasklets or workqueues; the kernel provides a thread scheduling model and priority management.
- The hardirq return IRQ_WAKE_THREAD atomic operation ensures no lost interrupts.
- The threaded handler can set different scheduling policies (e.g., SCHED_FIFO, SCHED_RR) and can also set CPU affinity.
- Reduce interrupt off time, improving system response speed.
10. Timer Interrupts and Kernel Time Management
10.1 The Importance of Timer Interrupts (Tick Interrupt)
Timer interrupts are the heartbeat of kernel time management. Prior to Linux 2.6.21, the system relied on a fixed-frequency periodic interrupt (tick period, typically 100Hz/250Hz/1000Hz). The main tasks of tick interrupts include:
- Updating wall time (jiffies, wall_time).
- Handling timer expiration (hrtimer, timer wheel).
- CPU time statistics (user_time, system_time).
- Task Scheduling Periodic Checks (scheduler_tick).
- RCU grace period checks.
10.2 Dynamic Tick
The overhead of a fixed tick is significant (CPU cannot enter deep sleep states). Linux introduced Dynamic Tick (config\_NO\_HZ) since 2.6.21. After enabling, the kernel dynamically adjusts the next tick based on the current pending task conditions, allowing the CPU to enter C-states deeper power saving when there are no pending tasks.
10.3 High-Resolution Timer
hrtimer (High-resolution Timer) is an independent module of the timer subsystem, which achieves nanosecond-level accuracy using hardware high-resolution events (HPET, TSC-deadline, LAPIC Timer), getting rid of the millisecond-level limitations of jiffies. hrtimer is internally implemented with red-black trees in the earliest expiration time and is used to implement nanosleep(), timerfd, and other high-precision timer interfaces.
11. Performance Monitoring and Tuning Tools
11.1 Viewing Interrupt Information (/proc/interrupts)
/proc/interrupts provides the count of each interrupt vector received on each CPU core and the corresponding device information. For example:
CPU0 CPU1 CPU2 CPU3
16: 12345 0 0 0 IO-APIC 16-fasteo enp0s3
17: 0 54321 0 0 IO-APIC 17-fasteo enp0s8
This file is extremely useful for diagnosing interrupt distribution unbalanced loads.
11.2 Softirq Statistics (/proc/softirqs)
/proc/softirqs provides the number of executions per softirq type per CPU:
CPU0 CPU1 CPU2 CPU3
HI: 1234 567 890 234
TIMER: 56789 3456 7890 1234
NET_TX: 123 456 789 234
NET_RX: 9876543 123456 78901 23456
11.3 Performance Analysis Tools
- perf stat -e irq:irq_handler_entry: Counts entry events for interrupt handlers.
- trace-cmd: Trace the complete interrupt processing path (through tracepoints such as
irq:irq_handler_entryandsoftirq:softirq_entry). - /proc/irq/IRQ#/spurious: View abnormal interrupt counts.
11.4 Interrupt Tuning Suggestions
- Interrupt Affinity Tuning: For high-throughput network cards, distribute different queue interrupts to different CPU cores using
smp_affinity. - Disable irqbalance: In high-performance computing scenarios, binding hardware interrupts manually is more reliable than automatic balancing.
- Adjusting NAPI Weight: Network card driver adjusts Polling Weight via
/sys/class/net/<if>/gro_flush_timeoutto balance throughput and latency. - workqueue Performance: Set WQ_HIGHPRI flag or use ordered workqueue for high-priority tasks.
12. Summary
The Linux kernel interrupt subsystem is the cornerstone that bridges hardware events and the kernel. From the trigger of hardware interrupts to the flow of driver handler functions passing to softirq/tasklet/workqueue, every design is committed to finding the optimal balance between "low latency" and "high throughput". Modern Linux recommends that device drivers prioritize using request_threaded_irq() because it naturally separates fast acknowledgment and slow processing, with a clear architecture and less programming burden.
Understanding the difference between interrupt context and process context, mastering the usage scenarios and constraint conditions of softirq, tasklet, and workqueue—will help system programmers write high-performance and high-reliability device drivers.
References
- Linux Kernel Source Code:
kernel/irq/,kernel/softirq.c,kernel/workqueue.c,arch/x86/kernel/irq.c - 《Understanding the Linux Kernel, 3rd Edition》— Bovet & Cesati
- 《Linux Device Drivers, 3rd Edition》— Corbet, Rubini, Kroah-Hartman
- Kernel Documentation:
Documentation/core-api/irq,Documentation/softirq.txt

发表评论 取消回复