Smart Farm Offline, Manual Mode & Fail-safe: Internet ล่ม Sensor เสีย แล้วระบบควรทำอะไร.
ออกแบบ Smart Farm ให้ fail gracefully เมื่อ Internet, Gateway, Sensor หรือ Server ล่ม ตั้งแต่ local autonomy, stale data, manual mode, command ownership, safe fallback, recovery และ protection boundary

Smart Farm ที่ดีต้องออกแบบตอน ระบบไม่สมบูรณ์ ตั้งแต่แรก ไม่ใช่ทดสอบเฉพาะตอน Sensor, Wi-Fi, Gateway, Server และ Cloud ทุกอย่าง online พร้อมกัน
คำถามหลักคือ:
ถ้าส่วนหนึ่งล่ม ระบบควรหยุด, ทำงานต่อแบบจำกัด, fallback หรือขอคนยืนยัน?
Failure ไม่ได้มีแบบเดียว
แยกอย่างน้อย:
- Sensor stale/offline
- field node offline
- gateway offline
- Internet down
- server/broker unavailable
- actuator feedback missing
- power interruption
- operator/network command conflict
คำว่า “offline” คำเดียวไม่พอสำหรับ control policy
Local Autonomy
ระบบสำคัญบางส่วนควรทำ local ได้ตาม bounded policy เช่น:
- tank low inhibit
- maximum pump runtime
- local irrigation schedule fallback
- manual start/stop ตาม authority
- local alarm indicator
ไม่จำเป็นว่าทุก intelligence ต้องอยู่ Cloud
Stale Data Policy
แต่ละ measurement ที่ใช้ควบคุมต้องกำหนด TTL/freshness expectation
fresh → normal policy
delayed → warning / limited use
stale → fallback or inhibit
missing → explicit unavailable state
อย่าค้าง last-known value แล้วทำเหมือนยังสด
Safe State ไม่จำเป็นต้อง OFF ทุกอย่าง
OFF อาจปลอดภัยในบางอุปกรณ์ แต่ไม่ใช่ทุก process
ต้องกำหนดตาม physical system เช่น:
- valve fail/open/closed behavior
- pump process consequence
- greenhouse ventilation
- water source/tank constraints
คำว่า fail-safe จึงต้องผูกกับ hazard/process ไม่ใช่ default software boolean
Manual Mode
Manual Mode ที่ใช้งานจริงควร:
- เข้า/ออก mode ได้ชัด
- แสดง owner ของ command
- log operator action
- ป้องกัน Auto เข้ามาแย่ง command
- ยังเคารพ protection/interlock ที่ต้องมี
Manual ไม่เท่ากับ “disable ทุก safety check”
Command Ownership
หนึ่ง actuator ไม่ควรถูกหลายระบบเขียน state แข่งกัน
ตัวอย่าง priority model ต้องกำหนดให้ชัด เช่น:
Maintenance lockout
> Local emergency/protection action
> Authorized manual mode
> Automatic policy
> Remote convenience command
ลำดับจริงขึ้นกับระบบ แต่หลักคือมี owner explicit
Internet Down
เมื่อ Internet ล่ม:
- field measurement อาจยังอ่านได้
- local node/gateway อาจยัง control ได้
- telemetry อาจ buffer
- remote dashboard unavailable
อย่าออกแบบให้ Internet down = สูญเสีย control ทั้งหมดโดยไม่จำเป็น
Gateway Restart
หลัง reboot:
- discover/read physical state
- invalidate expired command
- check sensor freshness
- restore only bounded local policy
- reconnect upstream
- publish current state/health
อย่า replay command เก่าทันทีเพราะ queue ยังมี message
Sensor Disagreement
ถ้ามีหลาย Sensor แล้วค่าไม่ตรง:
- flag disagreement
- isolate suspect input เมื่อ policy รองรับ
- fallback conservative
- require inspection ถ้า consequence สูง
อย่า average ความผิดปกติจนหายไปจาก dashboard
Feedback Missing
ถ้า command ถูกส่งแต่ไม่มี feedback ภายใน timeout:
requested ≠ confirmed
ระบบควรเข้าสถานะ fault/unknown ตาม design ไม่ใช่ assume success
Power Recovery
หลังไฟกลับ:
- output default state ต้องรู้
- pump/valve ไม่ควร start จาก retained command โดยไม่ verify
- stagger restart เมื่อจำเป็น
- gateway/server state ต้อง reconcile กับ physical state
Alarm During Offline
ถ้า cloud unavailable local system อาจต้องเก็บ alarm/event ไว้แล้ว sync ทีหลัง โดยรักษา original event time
Test Failure ก่อน Deploy
ควรมี test matrix เช่น:
| Test | คาดหวัง |
|---|---|
| ถอด Internet | local critical policy ยังอยู่ |
| ปิด broker/server | command path local ไม่พังทั้งหมด |
| Sensor stale | policy เข้า fallback ที่กำหนด |
| ถอด flow feedback | run ถูก bounded/alert |
| reboot gateway | ไม่ replay expired command |
| switch Manual | Auto ไม่แย่ง authority |
Protection Boundary
Automation ≠ Protection
Monitoring ≠ Safety
Connectivity ≠ Command Ownership
Software ไม่แทน breaker/fuse, overload, dry-run protection ที่จำเป็น, emergency isolation, pressure relief, equipment-native limits หรือ manufacturer requirement
อ่าน control coordination ที่ Pump & Valve Automation และ network behavior ที่ Remote Field Connectivity
ระบบ Smart Farm ที่ mature จึงไม่ได้พิสูจน์ตัวเองตอนทุกอย่าง online แต่พิสูจน์ตอน บางอย่างหายไปแล้วระบบยัง fail อย่างคาดเดาได้ และคนยังควบคุมสถานการณ์ได้



