
A technician seats a CPU into a motherboard socket during the assembly phase of semiconductor fabrication, where precision handling and controlled environments are critical to protecting delicate integrated circuits.
A fab floor is one of the most instrumented environments on the planet, and yet the gap between data collected and decisions made in time is where yield points and capacity commitments quietly disappear. When process drift goes undetected for even a few wafer lots, the cost compounds fast.
A single lithography excursion propagating through downstream steps before anyone catches it does not just damage those lots; it ties up equipment, disrupts WIP flow, and triggers expedite cycles that ripple across the schedule. Cleanroom events, tool availability drops, and cycle time excursions all operate the same way: invisible until they are expensive. The ten KPIs below are the ones that tell you what is actually happening inside your fab right now, not what happened during the last shift review.
Wafer Starts per Week
- Why it Matters: Wafer starts drive every downstream capacity and delivery commitment. Deviations from plan compound quickly across a fab’s long cycle time.
- What it Measures: The count of wafers entered into the production queue per week, compared against plan.
- What Happens if Missed: Shortfalls surface too late to recover within the same planning horizon, forcing schedule push-outs to customers.
- Formula: Actual Wafer Starts / Planned Wafer Starts x 100
- Indicator Type: Current. This KPI reflects live production entry rate and signals capacity alignment in the near term.
- Unit of Measure: Wafers per week (% of plan)
- Ideal Visualization(s): KPI trend with real-time alerts; bullet chart showing actual vs. plan
- Frequency: Daily rollup with real-time lot-entry tracking
- Data Required: Lot release timestamps, wafer count per lot, weekly production plan targets
- Pro Tip: Track by product family, not just total starts. Shortfalls on advanced node products are not equivalent to shortfalls on mature node products.
- Red Flag: Three or more consecutive days below 95% of plan usually signals an upstream material or equipment availability issue, not a scheduling artifact.
Fab Cycle Time
- Why it Matters: Cycle time is the primary lever on delivery performance and capacity utilization. Every extra day in queue is a day of customer lead time you cannot recover.
- What it Measures: The elapsed time from wafer start to completion of the last process step, compared against target cycle time.
- What Happens if Missed: Cycle time creep inflates WIP, reduces throughput efficiency, and breaks commit dates without any single obvious failure to point to.
- Formula: Actual Cycle Time / Target Cycle Time x 100
- Indicator Type: Lagging. Cycle time reflects the cumulative result of equipment availability, queue management, and process stability across all steps.
- Unit of Measure: Days (% of target)
- Ideal Visualization(s): KPI trend with real-time alerts; box plot for cycle time distribution by product layer
- Frequency: Updated per lot completion; rolling 7-day average
- Data Required: Lot start timestamp, lot completion timestamp, step-level queue wait times, target cycle time by product
- Pro Tip: Separate queue wait time from active process time. Most cycle time inflation hides in queue, not on tool.
- Red Flag: Cycle time rising while starts hold steady points to a bottleneck tool or a WIP imbalance building at a specific layer.
Gross Die Yield
- Why it Matters: Gross die yield is the most direct measure of fab process health and the primary driver of cost per good die. Small yield shifts at volume have large financial consequences.
- What it Measures: The percentage of functional dies produced per wafer relative to the theoretical die count at that reticle step.
- What Happens if Missed: Yield excursions that propagate across multiple lots before detection create write-off events that cannot be recovered within the affected quarter.
- Formula: (Good Die Count / Theoretical Die Count) x 100
- Indicator Type: Lagging. Yield integrates the quality of every process step a wafer has passed through.
- Unit of Measure: Percent (%)
- Ideal Visualization(s): SPC trend (control chart) with real-time alerts; KPI trend with real-time alerts by product and lot
- Frequency: Per lot at final electrical test; interim parametric yield at key in-line test steps
- Data Required: Good die count per wafer, theoretical die count per reticle, lot identifiers, process layer, product code
- Pro Tip: Watch yield by process layer, not just final test. Parametric yield at intermediate steps often predicts final yield weeks earlier.
- Red Flag: A yield drop confined to a specific layer or tool group almost always points to a process excursion or equipment maintenance issue, not a random event.
Critical Tool Availability
- Why it Matters: Availability of bottleneck tools drives fab throughput more than any scheduling decision you can make. A lithography stepper or etch chamber down at the wrong time cascades across the entire floor.
- What it Measures: The percentage of scheduled production time that a critical tool is available and running, excluding planned maintenance windows.
- What Happens if Missed: Unplanned downtime on a bottleneck tool creates WIP starvation at downstream steps and ripples through delivery schedules within hours.
- Formula: (Scheduled Time – Unplanned Downtime) / Scheduled Time x 100
- Indicator Type: Current. This KPI reflects live equipment state and responds immediately to faults, alarms, and maintenance events.
- Unit of Measure: Percent (%)
- Ideal Visualization(s): KPI blocks for each critical tool; Pareto chart when ranking tools by unplanned downtime contribution; KPI trend with real-time alerts
- Frequency: Real-time, updated per tool state change
- Data Required: Tool scheduled production time, unplanned downtime events with timestamps, tool state (running, idle, down, maintenance), tool classification
- Pro Tip: Track availability separately for bottleneck tools vs. non-bottleneck tools. Aggregate fab availability numbers mask the equipment that actually controls throughput.
- Red Flag: Availability trending down gradually over two to three weeks often signals a tool that needs preventive maintenance before it causes an unplanned stop.
Defect Density (D0)
- Why it Matters: Defect density is the leading process health signal at every layer. It predicts yield before electrical test data is available and identifies contamination events early enough to contain them.
- What it Measures: The number of particle or pattern defects per square centimeter on the wafer surface at a given process step.
- What Happens if Missed: A defect density excursion that is not caught at the inspection step will be confirmed only at final yield, by which point multiple additional lots have passed through the same compromised process.
- Formula: Total Defects / Inspected Wafer Area (cm²)
- Indicator Type: Leading. D0 at in-line inspection predicts final electrical yield and flags equipment or process issues before they fully express as failures.
- Unit of Measure: Defects per cm²
- Ideal Visualization(s): SPC trend (control chart) with real-time alerts; KPI trend with real-time alerts by inspection layer and tool
- Frequency: Per lot at each scheduled inspection point
- Data Required: Defect count per wafer, inspected area per wafer, inspection step, lot identifier, tool used at prior process step
- Pro Tip: Correlate D0 spikes with specific tool IDs at the preceding process step. Most contamination events have a single tool source that shows up clearly in the data.
- Red Flag: D0 rising on a single inspection layer while adjacent layers hold steady is a strong signal of a localized tool or chamber problem, not a systemic process drift.
Cleanroom Particle Count
- Why it Matters: The cleanroom environment is a shared resource. A particle event anywhere in the fab threatens yield across all product types running on the floor simultaneously.
- What it Measures: Airborne particle concentration in the production environment, measured at defined monitoring points by particle size class.
- What Happens if Missed: Particle excursions that are not caught in real time expose every wafer on the floor to potential contamination, with yield impact that only becomes visible at downstream inspection or test.
- Formula: N/A
- Indicator Type: Leading. Particle count in the environment directly predicts defect density risk on wafers before any in-process inspection occurs.
- Unit of Measure: Particles per cubic meter (by size class)
- Ideal Visualization(s): KPI trend with real-time alerts; Group Map showing particle count status by cleanroom zone
- Frequency: Continuous real-time monitoring, updated per sensor sample interval
- Data Required: Particle counts by size class, sensor location identifiers, cleanroom zone mapping, ISO class thresholds per zone
- Pro Tip: Map particle sensor alerts to concurrent wafer lot locations on the floor. This lets you immediately identify which lots were at risk during an excursion.
- Red Flag: A particle spike isolated to one zone with no corresponding increase in adjacent zones almost always points to a local equipment or process event, such as a robot fault or filter bypass.
Photolithography CD Uniformity (Within-Wafer)
- Why it Matters: Critical dimension variation at the lithography step propagates directly into device performance and yield. CD control is the foundation of every node’s process window.
- What it Measures: The standard deviation or range of patterned feature dimensions across the wafer surface at a given lithography step.
- What Happens if Missed: CD excursions that exceed the process window cause parametric failures at electrical tests, often appearing as yield loss that is difficult to trace back to the original source step.
- Formula: N/A
- Indicator Type: Leading. CD uniformity at the patterning step predicts parametric yield and device performance before any downstream electrical test.
- Unit of Measure: Nanometers (nm), expressed as 3σ
- Ideal Visualization(s): SPC trend (control chart) with real-time alerts; KPI trend with real-time alerts by stepper tool and lot
- Frequency: Per lot at lithography inspection; real-time overlay from inline metrology tools
- Data Required: CD measurements by wafer location, lot identifier, stepper tool ID, reticle ID, process layer, target CD value
- Pro Tip: Separate within-wafer uniformity from lot-to-lot variation. Systematic within-wafer patterns often point to illumination or focus issues on a specific tool.
- Red Flag: CD drift that tracks a specific reticle across multiple tools rules out the stepper as the source and implicates the mask itself.
WIP (Work in Progress) Lot Velocity
- Why it Matters: WIP balance drives cycle time efficiency. Lots stacking at a specific process layer signal a bottleneck or scheduling problem that will translate into missed delivery dates within days.
- What it Measures: The average number of process steps completed per lot per day, compared against plan, across the active WIP population.
- What Happens if Missed: WIP accumulation at bottleneck steps inflates cycle time silently until it is visible in delivery performance, at which point the recovery window is already narrow.
- Formula: (Steps Completed per Lot) / (Elapsed Days in Fab) vs. Plan
- Indicator Type: Current. Lot velocity reflects the real-time flow rate of active WIP and responds immediately to tool downtime and queue imbalances.
- Unit of Measure: Steps per day (% of plan)
- Ideal Visualization(s): KPI trend with real-time alerts; Group rollup bars by process area or layer group; Pareto chart when ranking layers by WIP accumulation
- Frequency: Updated continuously as lots complete steps; hourly summary
- Data Required: Step completion timestamps, lot identifiers, planned step sequence, current step for each active lot
- Pro Tip: Track velocity by process area, not just total fab. A velocity problem in lithography looks completely different from one in etch or CMP, and the response is different too.
- Red Flag: Velocity holding on track overall while WIP accumulates at one specific layer is the classic signature of a hidden bottleneck that aggregate metrics will not surface.
Etch Rate Uniformity
- Why it Matters: Etch rate variation within a wafer or between chambers drives CD deviation, profile changes, and selectivity failures that directly affect device characteristics and yield.
- What it Measures: The uniformity of material removal rate across the wafer surface during dry or wet etch processes, expressed as a percentage variation from the mean.
- What Happens if Missed: Chamber drift that is not caught before production lots are processed creates systematic defects across every wafer run on that tool until the issue is identified and corrected.
- Formula: (Max Etch Rate – Min Etch Rate) / (2 x Mean Etch Rate) x 100
- Indicator Type: Leading. Etch uniformity measured on monitor wafers or via in-situ endpoints predicts process stability before product wafers are exposed to the same conditions.
- Unit of Measure: Percent (%) non-uniformity
- Ideal Visualization(s): SPC trend (control chart) with real-time alerts; Pareto chart when ranking etch chambers by non-uniformity contribution; KPI trend with real-time alerts by chamber ID
- Frequency: Per lot or per chamber qualification run; real-time where endpoint detection is available
- Data Required: Etch depth measurements by wafer location, mean etch rate, chamber identifier, process recipe, lot identifier, monitor wafer results
- Pro Tip: Run chamber-matching analysis regularly across your etch tool set. Systematic differences between chambers cause within-lot yield variation that is easy to miss when you look at averages.
- Red Flag: Etch rate non-uniformity trending upward on a single chamber while sister chambers hold control is the clearest signal of a chamber component issue such as a worn electrode or blocked gas inlet.
Overall Equipment Effectiveness (OEE) for Critical Tools
- Why it Matters: OEE for bottleneck tools ties together availability, throughput rate, and quality in a single number. It is the most complete view of whether a tool is delivering its full production capacity.
- What it Measures: The product of tool availability, performance (actual vs. rated throughput), and quality (lots processed without excursion) for a defined set of critical tools.
- What Happens if Missed: OEE drift that is not tracked in real time allows compounding inefficiencies across availability, speed loss, and quality losses to accumulate before any single metric triggers a response.
- Formula: Availability (%) x Performance (%) x Quality (%)
- Indicator Type: Current. OEE integrates live tool state data and responds to changes in any of its three components in real time.
- Unit of Measure: Percent (%)
- Ideal Visualization(s): KPI blocks by critical tool; KPI trend with real-time alerts; Pareto chart when ranking tools by OEE loss contribution
- Frequency: Real-time, updated per tool state change and lot completion
- Data Required: Tool scheduled time, unplanned downtime, actual run rate, rated run rate, lots processed, lots with quality excursion, tool classification
- Pro Tip: Decompose OEE losses into their three components when diagnosing performance problems. Most tools lose more to speed and quality losses than to downtime, but downtime gets all the attention.
- Red Flag: OEE trending down on a tool where availability is stable almost always means performance or quality losses are accumulating, often from process drift or recipe creep that maintenance checks will not catch.
Why Real-Time Visibility Matters
Semiconductor fabs operate at process tolerances measured in nanometers, across hundreds of sequential steps, with cycle times stretching across weeks. In that environment, a process excursion at layer 10 does not announce itself at layer 10. It shows up at electrical tests, by which point the affected lots have traveled through dozens more steps and consumed capacity that cannot be returned. Real-time KPIs at every critical layer close that gap before it becomes a write-off event.
The same logic applies to equipment availability, WIP flow, and cleanroom integrity. Each of these operates on a time scale where a few hours of missed visibility translates directly into schedule impact that no expedite process can fully recover. The KPIs above are not reporting metrics; they are the operational controls that keep a fab running at plan instead of reacting to the last thing that went wrong.
How Transpara Can Help
If real-time operational visibility is a challenge you’re facing, you’re not alone. At Transpara, we help teams like yours gain clarity from complex systems without the need to centralize or overhaul your data stack.
Learn more about Transpara
Browse our documentation
Contact us