You've seen the demos. A camera watches parts go by, red boxes appear around defects, everyone claps. Clean. Impressive. Completely silent about the 47 decisions you need to make between "that looks cool" and "it's running on our line."
This is the guide to those 47 decisions.
Computer vision in manufacturing uses cameras and deep learning to automate tasks that currently require human eyes:
Where it struggles:
Four layers. Get any of them wrong and the system underperforms.
Industrial cameras. Not consumer webcams, not IP security cameras. The choice depends on what you're inspecting:
Match resolution to the smallest defect you need to catch. Higher resolution means slower processing and higher cost with no detection benefit if your defects are visible at lower resolution.
This is where most first deployments go wrong.
Defects invisible under ambient factory lighting become obvious under structured illumination. The type of lighting changes what's detectable:
Budget 30-40% of your hardware spend on lighting. If that sounds high, consider the alternative: a perfectly good camera and model that can't see the defects you need to catch.
Inference happens on the factory floor, not in the cloud. Response time requirements are milliseconds, not seconds.
At 100 units/minute, you have 600ms per inference. Most platforms handle that. At 1,000 units/minute, you need sub-60ms — that narrows your options to Jetson AGX or dedicated GPU.
The part vendors don't demo. Trigger a camera capture, run inference, and execute an action (reject, alert, log) — all within cycle time. This means PLC or SCADA integration via OPC UA or direct I/O.
A model that detects defects perfectly but can't trigger a reject mechanism within your cycle time is a science project.
The biggest misconception about computer vision in manufacturing is that you need tens of thousands of labeled images. You don't.
The bottleneck is image quality, not quantity. Your dataset needs to cover the full range of production conditions: different materials, lighting variation, product orientations, defect severities. A thousand images from identical conditions teaches the model less than 500 from varied conditions.
| Phase | Duration | What Happens |
|---|---|---|
| Assessment | 2-4 weeks | Define requirements, select camera/lighting, identify integration points |
| Hardware install | 2-4 weeks | Mount cameras, lighting, edge compute. Connect to PLC/SCADA |
| Data collection | 2-4 weeks | Capture images across operating conditions. Label defect types |
| Model training | 2-4 weeks | Train initial models, validate against test set |
| Parallel operation | 4-8 weeks | Run alongside existing inspection. Tune thresholds. Build operator trust |
| Cutover | 1-2 weeks | Switch to primary. Establish monitoring and retraining cadence |
Total: 3-6 months for a single line. Subsequent lines are faster — the data pipeline and model architecture are reusable.
Realistic range for a single inspection station:
| Component | Range |
|---|---|
| Cameras + lighting | $5K-$30K |
| Edge computing | $2K-$10K |
| Integration & installation | $10K-$30K |
| Software (vendor platform or custom) | $15K-$50K |
| Training & commissioning | $5K-$15K |
| Total | $40K-$135K |
Most manufacturers report payback within 6-12 months from reduced scrap, rework, warranty costs, and labor reallocation.
Rank your inspection stations on three criteria:
Cost of escaped defects. Where do missed defects hurt most? Warranty claims, recalls, downstream rework, customer complaints?
Current inspection reliability. Where are you most dependent on human consistency? High-volume lines where sampling is the only option? Shifts with known quality dips?
Visual detectability. Can a camera see the defect? Internal voids and material composition issues need different sensing. Computer vision works when there's a visual signature.
Deploy where all three scores are high. One line, one defect type, proven ROI, then expand.
Skimping on lighting. Cameras and compute get the budget. Lighting gets the leftovers. Flip that priority. Lighting determines whether defects are visible at all.
Trying to catch everything on day one. Start with 2-3 defect types. Nail them. Expand later. Trying to detect 15 defect types at launch is how projects stall.
Ignoring integration. If the system can't trigger a physical reject within cycle time, it's not a production tool.
Expecting perfection immediately. Threshold tuning takes 4-8 weeks of parallel operation. Plan for it. Operators need to see the system's calls validated before they trust it.