- 技术排查通话的第一个问题应该是什么?What should be the first question in a troubleshooting call?
- 如何区分模型误报、成像变化和接口延迟?How can you separate a model false alert, imaging change and interface delay?
- 什么情况下可以说 root cause confirmed?When can you say the root cause is confirmed?
- 临时关闭一个检测规则时需要记录什么?What must be recorded when a check is temporarily disabled?
处理现场技术问题Handling an On-site Technical Issue
练习按现象、影响、时间线、隔离测试、临时措施和根因验证的顺序,处理相机画面、网络、误报和设备联动问题。Practise handling camera-image, network, false-alert and device-integration issues through symptoms, impact, timeline, isolation tests, containment and root-cause verification.
MIN / SESSION
TO
VISION
Handling an On-site Technical Issue
本期学习内容已经发布,真实 MP3 尚待完成 TTS 配置后生成。当前不会显示失效或虚假的播放状态。The learning material is published. The real MP3 will be generated after a TTS provider is configured; no broken or simulated playback is shown.
- 在猜测原因前准确描述现象和影响范围。Describe the symptom and impact before guessing at a cause.
- 用一次只改变一个条件的方式设计隔离测试。Design isolation tests that change one condition at a time.
- 区分临时恢复措施、初步发现和已验证根因。Distinguish containment, preliminary findings and a verified root cause.
Build a clear answer for the real situation.
一个已在验证阶段运行的工位突然出现间歇性 NG 提示。你需要与客户工程师远程协作,在不干扰生产的前提下定位问题。A station in validation begins showing intermittent NG prompts. You must work remotely with the customer engineer to locate the issue without disrupting production.
依次询问开始时间、频率、影响范围、最近变化和可保存证据。Ask about start time, frequency, affected scope, recent changes and available evidence in order.
90 SEC- 先描述 symptom,不说 root cause。Describe the symptom without claiming a root cause.
- 询问最近变化。Ask what changed recently.
- 提出只改变一个变量的测试。Propose a one-variable test.
使用 confirmed、indication、remains under investigation 区分证据强度。Use ‘confirmed’, ‘indication’ and ‘remains under investigation’ to distinguish evidence strength.
60 SEC- 说明已恢复的部分。State what has recovered.
- 说明尚未解决的部分。State what remains unresolved.
- 给出负责人和下一复核点。Give an owner and next review point.
Read both roles, then repeat one role aloud.
Ben, the station has started showing NG prompts even when the operator completes the step correctly.
Ben,即使操作员正确完成步骤,工位也开始显示 NG 提示。Let us separate the symptom from the possible cause. When did it start, how often does it occur and which product variants are affected?
我们先把现象与可能原因分开。它何时开始、多久发生一次、影响哪些产品变型?It began after the lunch break. It is intermittent and appears more often on the darker variant.
午休后开始出现,具有间歇性,而且在深色变型上更常见。Has anything changed in the station environment, camera position, lighting, model version or interface configuration since the last normal run?
自上次正常运行后,工位环境、相机位置、光照、模型版本或接口配置是否有变化?No software was deployed. Maintenance cleaned the work area, and the adjustable lamp may have been moved.
没有部署软件。维护人员清洁了作业区域,可调灯可能被移动过。Please save one normal frame and one false NG frame from the same variant. Do not move anything yet; I want to compare brightness, reflection and target position first.
请保存同一变型的一帧正常画面和一帧错误 NG 画面。先不要移动任何东西;我想先比较亮度、反光和目标位置。The false NG frame looks brighter on the left edge, and some detail in the target area is washed out.
错误 NG 画面左侧更亮,目标区域的一些细节出现过曝。That is a useful clue, but not yet a confirmed cause. Can you check the lamp angle against the approved setup photo?
这是一个有用线索,但还不是已确认原因。你能把灯具角度与获批设置照片对照吗?The lamp is pointing lower than in the photo. Should I return it to the marked position?
灯比照片里指向更低。我应该把它恢复到标记位置吗?Yes, if it is safe and permitted. Change only the lamp position, then run the approved check samples again. That will help us isolate the variable.
可以,前提是安全且获准。只改变灯的位置,然后再次运行获批检查样本,这有助于隔离变量。After the adjustment, the normal samples pass again, including the darker variant.
调整后,正常样本再次通过,包括深色变型。Good. Now repeat a small controlled set and monitor both the vision result and the tool signal timestamp. We need to rule out a second issue before returning to normal validation.
很好。现在重复一小组受控样本,同时监控视觉结果和工具信号时间戳。恢复正常验证前,我们需要排除第二个问题。The images are stable, but one result arrived after the workflow timeout.
图像稳定了,但有一个结果在工作流超时后才到达。Then we may have two separate symptoms. Please export the local event timestamps without any sensitive production fields. I will compare capture, inference and workflow receipt times.
那么可能存在两个独立现象。请导出本地事件时间戳,不要包含敏感生产字段。我会比较采集、推理和流程接收时间。For now, should we keep the station in observation mode?
目前我们是否应让工位保持观察模式?Yes. Record that as a temporary containment measure, with an owner and review time. We can say the lighting issue is reproduced and corrected, but the delayed result remains under investigation.
是的。把它记录为临时控制措施,注明负责人和复核时间。我们可以说光照问题已复现并纠正,但结果延迟仍在调查中。防止排查一开始就锁定错误方向。Avoid locking onto an unverified cause too early.
Let us separate the symptom from the possible cause.
我们先把现象与可能原因分开。描述并非每次发生的问题。Describe a problem that does not occur every time.
The false alert is intermittent.
错误报警是间歇性的。控制结论强度。Control the strength of a conclusion.
The reflection is a useful clue, but not yet a confirmed cause.
反光是有用线索,但还不是已确认原因。说明一次只改变一个条件的测试目的。Explain why only one test condition changes at a time.
Change only the lamp angle to isolate the variable.
只改变灯具角度以隔离变量。说明测试用于排除另一种可能。State that a test excludes another possibility.
We need to rule out an interface delay.
我们需要排除接口延迟。透明说明未解决问题。Communicate an unresolved issue transparently.
The delayed result remains under investigation.
结果延迟仍在调查中。两个问句保持同一节奏。Keep the same rhythm in both questions.
在 clue 后停顿,重读 not yet。Pause after ‘clue’ and stress ‘not yet’.
重读 only one。Stress ‘only one’.
rule out 连读。Link ‘rule out’.
使用平稳、不防御的语气。Use a steady, non-defensive tone.
先报告可观察现象,不把问题笼统归因于 AI。Report the observable symptom instead of broadly blaming ‘AI’.
视觉差异是线索,需要通过复现实验确认。A visual difference is a clue that requires a repeat test.
同时改变多个变量会破坏因果判断。Changing several variables destroys causal clarity.
区分领先假设与已确认原因。It distinguishes a leading hypothesis from a confirmed cause.
准确拆分已解决和未解决事项。It separates resolved and unresolved items accurately.
Listen for purpose, detail and next actions.
现场工程师和远程支持人员共同调查两个交织问题:灯具位置变化导致图像过曝,以及网络抖动造成结果延迟。An on-site engineer and remote specialist investigate two overlapping issues: image overexposure caused by a moved lamp and delayed results caused by network instability.
- 哪条证据支持光照假设?Which evidence supports the lighting hypothesis?
- 为什么光照恢复后调查仍未结束?Why does the investigation continue after the lighting is restored?
- 双方如何保护生产并保留排查证据?How do they protect production and preserve troubleshooting evidence?
写下图像问题和时序问题各自的现象、证据和当前状态。Write the symptom, evidence and current status for the image issue and timing issue.
现象与状态Symptoms and status按发现、测试、恢复、再次异常、临时措施的顺序记录事件。Record events in the order of detection, test, recovery, second symptom and containment.
时间线与证据强度Timeline and evidence strength阅读、聆听,并大声说出来Read, listen and speak aloud
Jack, I have paused the validation run and switched the station to observation mode. Production can continue, but the AI result is not controlling any action.
Good containment. Please give me the exact symptom and the last time the station ran normally.
After lunch, correct operations began receiving occasional NG prompts. The morning run was normal. No model or workflow version changed between the two runs.
Have you compared the configuration checksum and camera status, rather than relying only on the deployment record?
Yes. The approved versions and camera settings match. However, the image histogram is brighter than the morning reference.
That points us towards the physical environment. Check the lamp position, any new reflective object and whether the camera bracket has moved.
The camera bracket marks are aligned. A cleaning trolley was removed, and the task lamp is below its marked angle.
Return only the lamp to its approved mark, take a reference frame and run the controlled sample set. Keep the exposure settings unchanged.
The left-edge glare has disappeared. All approved normal samples now produce the expected vision state.
We have reproduced and corrected the image symptom. Save the before-and-after frames and record the lamp position as the verified cause for that symptom.
During the repeat run, I noticed one workflow timeout even though the image result was correct.
That should be tracked separately. What are the capture, inference-complete and client-receipt timestamps for that event?
Capture and inference timing are within the normal range. The delay appears between the edge unit sending the result and the client receiving it.
Check the local link status and packet-loss counter for that period. Please do not restart the switch yet, because that would remove useful evidence.
The edge link stayed up, and its packet-loss counter did not increase. The client host shows a short scheduling delay at the same time, so the current evidence does not prove a network fault.
Then the network hypothesis is weaker than we first thought. Keep it unconfirmed, and ask the system owner to compare host load, message-queue timing and security events before we change any timeout.
Should I increase the timeout now?
Not during production and not without reviewing the process limit. Keep observation mode active, preserve the logs and schedule a controlled network test with the network owner.
I will document two separate incidents: lighting position corrected, and intermittent result delivery under investigation.
Exactly. Add the temporary operating mode, owners, evidence links and next review time. That gives the team a truthful status without mixing the two causes.
The system owner has joined the call. The client host was running a scheduled security scan at the time of the delayed result. Could that explain the scheduling delay?
It is a plausible hypothesis, but timing overlap alone is not proof. We should compare processor load, storage activity, message-queue latency and several unaffected events from the same scan period.
The host log shows a short processor peak. Most events still arrived within the workflow limit, while one message waited longer in the client queue.
That narrows the location of the delay, but we still need a controlled reproduction. Ask the system owner whether the scan can be run in an approved test window while we replay a non-production event sequence.
The test window is available after the current shift. We can use the simulator, so no real station command will be issued.
Good. Record the host baseline first, run the event sequence without the scan, then repeat it with the approved scan. Keep the workflow version, queue settings and simulator input unchanged.
If the delay appears again, can we call the scan the root cause?
Only if the behaviour is repeatable, the comparison shows the queue delay, and alternative causes are reasonably ruled out. We should also confirm whether the host resource policy meets the deployment requirement.
Would increasing the workflow timeout be an acceptable containment measure?
Possibly, but only after the process owner confirms the maximum useful response time. A larger timeout might hide a performance problem or allow the operator to move ahead before the prompt arrives.
For the next production period, we will keep observation mode and prevent the security scan from overlapping with validation, subject to IT approval.
That is a reasonable temporary plan. Make it time-limited and record that it reduces exposure but does not yet prove the technical cause. The controlled test still needs to be completed.
Should we add monitoring so that a similar delay is easier to diagnose next time?
Yes. We should retain bounded metrics for inference time, transport time, client queue time and workflow evaluation time, with clock synchronisation checked. That makes each stage visible without retaining unnecessary production content.
I will update the incident record with the confirmed lighting cause, the unconfirmed host-load hypothesis, the containment plan and the approved test procedure.
Please also record the recovery criteria: stable images, repeatable message timing within the approved limit, completed log review and joint approval to leave observation mode. Then the closure decision will be evidence-based.
Understood. Until those criteria are met, the status will remain partially recovered rather than resolved. I will brief the next shift on the operating mode and escalation route.
That wording is accurate. It protects production, preserves the distinction between the two symptoms and gives the next team a clear testable route to full recovery.
Turn the dialogue into language you can use.
a temporary action that limits impact while investigation continues
Observation mode is the current containment measure.a value used to verify that data or configuration has not changed
The configuration checksum matches the approved version.a chart showing the distribution of image brightness values
The histogram shows that the image became brighter.strong reflected light that makes an image hard to inspect
The moved lamp created glare near the target.to make the same symptom happen again under controlled conditions
The team reproduced the image symptom.a workflow event triggered when a result does not arrive within a defined time
A delayed result caused a workflow timeout.data packets that fail to reach their destination
The counter showed a small increase in packet loss.a brief network link disconnection and reconnection
A link flap occurred during the delayed event.the defined way a system resends data after a failure
The controlled test will verify retry behaviour.to keep logs and records unchanged for investigation
Do not restart the switch until the evidence is preserved.把排查聚焦到可观察信息。Focus troubleshooting on observable facts.
Please give me the exact symptom and the last normal time.交叉验证记录。Cross-check a record.
Have you compared the checksum rather than relying only on the deployment log?设计单变量测试。Design a one-variable test.
Return only the lamp and keep exposure unchanged.拆分可能不相关的现象。Separate potentially unrelated symptoms.
The workflow timeout should be tracked separately.说明证据方向但保留根因确认边界。State what the evidence indicates while preserving root-cause ownership.
That evidence points to a host-side delay, but the system owner should confirm the cause.明确拒绝高风险即时变更。Reject a risky immediate change clearly.
Not during production and not without reviewing the process limit.1. 现场首先采取了什么临时措施?What containment was applied first?
暂停验证并切换到观察模式,生产继续但 AI 不控制动作。The validation was paused and the station switched to observation mode, allowing production to continue without AI control.
2. 什么证据首先指向物理环境变化?What first pointed to a physical-environment change?
- 模型版本改变A model-version change
- 图像直方图比参考更亮The image histogram was brighter than the reference
- PLC 写入失败PLC write-back failed
图像直方图比上午参考更亮。The image histogram was brighter than the morning reference.
3. 团队如何确认灯具位置是图像问题原因?How did the team confirm the lamp position caused the image issue?
只恢复灯具位置、保持曝光不变并重跑受控样本,眩光消失且正常样本通过。They restored only the lamp, kept exposure unchanged and reran controlled samples; the glare disappeared and normal samples passed.
4. 第二个问题发生在哪个时段?Where did the second issue occur in the timing chain?
发生在边缘单元发送结果与客户端接收结果之间。Between the edge unit sending the result and the client receiving it.
5. 为什么不立即重启交换机?Why do they avoid restarting the switch immediately?
- 交换机太重The switch is too heavy
- 重启会清除有用证据A restart would remove useful evidence
- 没有电源There is no power
因为重启可能清除排查所需证据。Because a restart could remove useful troubleshooting evidence.
6. 最终状态如何区分两个事件?How does the final status distinguish the two incidents?
灯具位置问题已纠正;间歇性结果传输仍在调查,并保持观察模式。The lamp-position issue is corrected; intermittent result delivery remains under investigation, with observation mode retained.
exact symptom、rather than、points us towards。Phrasing of ‘exact symptom’, ‘rather than’ and ‘points us towards’.
Good containment. Please give me the exact symptom and the last time the station ran normally. Have you compared the configuration checksum and camera status, rather than relying only on the deployment record? That points us towards the physical environment. Check the lamp position, any new reflective object and whether the camera bracket has moved.3 REPEATS
only、verified cause、tracked separately。Stress ‘only’, ‘verified cause’ and ‘tracked separately’.
Return only the lamp to its approved mark, take a reference frame and run the controlled sample set. Keep the exposure settings unchanged. We have reproduced and corrected the image symptom. Save the before-and-after frames and record the lamp position as the verified cause for that symptom. That should be tracked separately. What are the capture, inference-complete and client-receipt timestamps for that event?3 REPEATS
否定句的坚定语气与 but 的转折。Firm negatives and contrast with ‘but’.
Check the local link status and packet-loss counter for that period. Please do not restart the switch yet, because that would remove useful evidence. Then the network hypothesis is weaker than we first thought. Keep it unconfirmed, and ask the system owner to compare host load, message-queue timing and security events before we change any timeout. Not during production and not without reviewing the process limit. Keep observation mode active, preserve the logs and schedule a controlled network test with the network owner. Exactly. Add the temporary operating mode, owners, evidence links and next review time. That gives the team a truthful status without mixing the two causes.2 REPEATS
- 间歇性 NG 与版本排除Intermittent NG and version checks
- 灯具单变量测试与图像恢复Lamp isolation test and image recovery
- 结果延迟与网络证据Result delay and network evidence
- 观察模式、负责人和下一测试Observation mode, owners and next test
使用 symptom、hypothesis、verified cause 和 under investigation 四个不同证据等级。Use four evidence levels correctly: ‘symptom’, ‘hypothesis’, ‘verified cause’ and ‘under investigation’.
100 SEC RETELLING听 · 读 · 跟读 · 表达Listen · Read · Shadow · Speak
按顺序完成训练,状态只保存在当前设备,不上传任何个人数据。Complete the sequence in order. Progress stays on this device and is never uploaded.