Vision Intelligence Opens an Industrial “One-to-Many” Era | Interview with Lingbang Chairman Cui Zhongwei
Beijing Lingbang Intelligent Equipment Co., Ltd.
Chairman Cui Zhongwei
AI is reshaping industrial vision. Large-model visual agents are breaking the old “one station, one industrial PC” pattern. Lingbang’s VisionAgent has moved from 2.0 to 3.0, with a leap in speed and in how many tasks one unit can carry. The centralized edge plan — “one-to-many, and very many” — is a new way to do industrial inspection. In this conversation, chairman Cui Zhongwei walks through the core of VisionAgent 3.0, what makes the product different, and a live case: how the next generation of visual agents cuts cost and rebuilds the inspection layout.
Watch the interview
Mr. Cui, we already know the visual agent. AI is moving fast, and VisionAgent has gone from 2.0 to 3.0. What is new compared with the last generation?
Q1
We released the agent in October 2024. The first major break from 2.0 to 3.0 is speed. Inference on 2.0 was 40 FPS; 3.0 is already at 240 FPS. The agent’s biggest strength is few-shot, cross-domain generalization — but it cannot be slow. Only if it is fast enough does the commercial value of “one-to-many, and very many” hold up.
You mentioned “one-to-many, and very many.” What does that actually mean?
Q2
Because the agent is fast, it can pull a lot. There are two sides. First, along one line it can pull different tasks vertically: industrial inspection, process monitoring, assembly correctness — all sent to the same agent. Second, it can pull many stations and many lines at once. One device carrying all of that is “one-to-many, and very many.”
How can one VisionAgent carry so many line stations and tasks at the same time?
Q3
First, our own vision foundation model is few-shot and cross-domain. A small target of six pixels can still be recognized well. Second, every operator is GPU-native, so it stays fast. With those two, we can pull completely different tasks on a line, and do one-to-many at a large scale. We just talked about the edge agent; there is also the local agent. The local agent embeds a controller: 16 cameras, 31 lights, plus photoelectric triggers and motion control. That product is mainly so equipment makers can integrate a full system.
People care about real deployment. Can you give a concrete one-to-many case?
Q4
On disposable-tray inspection we did one-to-many at scale. The agent pulls 16 lines, five cameras on each line — 80 4K cameras in all. The camera signals come in over the network. From those 80 streams the agent judges NG or OK and writes the result back to the on-site PLC. A robot places each product on a conveyor; it passes five cameras; the agent returns the signal to the PLC; the PLC tells the robot how to sort. Throughput is 200+; the takt is a 20-second matrix, so five camera images enter the agent within 20 seconds. Divide image size by total throughput and you get the same picture: 16 lines, 80 4K cameras, edge-deployed one-to-many, and a better cost curve.
The traditional setup is one industrial PC per station. What is the advantage of VisionAgent’s centralized deployment?
Q5
One PC per station often means not enough compute, or leftover compute — the resource is hard to use fully. With one-to-many at the edge, we can fill the compute. That comes from the foundation model and the agent: they are general, easy to use, and fast to deliver. That is why this way of working uses the resource well and has a strong cost-performance ratio.