Momentrack — Real-Time Tracking of Emergency-Response Resources for Canada’s Department of National Defence

In 2019, my company Momentaj won a federal R&D contract from Canada’s Department of National Defence (DND) through Innovative Solutions Canada. It was meant to answer a question first responders had been asking for years. During a major incident, police, fire and paramedics from several agencies work the same scene. How does anyone know where every truck, kit and responder is, and whether each one is available?

I led the project and designed the solution. In six months, a team of nine built Momentrack. It had low-power tags, a vehicle gateway, and a private cloud that put every tracked resource on a live incident map. We delivered every milestone on schedule and moved the concept from technology readiness level (TRL) 3 to TRL 4.

At a glance

  • My role: Project Lead & Solution Architect, through Momentaj
  • Years: 2019–2020
  • Client: Department of National Defence, through Defence Research and Development Canada’s Centre for Security Science
  • Program: Innovative Solutions Canada (ISC), Challenge 29: Logistics and Resource Management of Emergency Response Assets
  • Selection: Momentaj was one of six companies awarded a Phase 1 contract. DND’s evaluators scored our proposal 96.5 out of 110, with full marks on eight of the twelve scored criteria, including innovation.
  • Delivery: December 2019 to May 2020, in five milestones, all on schedule
  • Team: nine people across hardware, firmware, cloud, UX and business analysis
  • What we built: four device types, three radio technologies (Bluetooth Low Energy, LoRa and LTE-M), a private cloud of seven services, and a live incident map
  • Result: TRL 3 → TRL 4 at both system and subsystem level, proven in ten tests, including a road test in a moving vehicle

TRL is the 1–9 technology readiness scale the Government of Canada uses. TRL 4 means the components have been proven to work together in a lab.


The challenge

DND’s Centre for Security Science runs the Canadian Safety and Security Program with Public Safety Canada, and its users are civilian first responders. In December 2018 it posted a challenge through ISC. It wanted real-time decision support for the people in command during critical incidents, when municipal, provincial and federal agencies, including the RCMP and DND, work together.

At the time, those teams tracked their resources with everything from spreadsheets to specialized software, and nothing gave them one shared picture. The challenge cited Project Responder 5, a US Department of Homeland Security study. In that study, responders ranked “the ability to know, in real time, the availability, location and status of all resources” among their most important capability needs. Canadian responders had described the same gap.

DND set six must-haves. A solution had to:

  • let responders request resources from the field and track each request
  • keep an inventory of what front-line vehicles carry
  • build one repository of every resource available for an incident
  • show critical resources on an incident map
  • use standardized icons
  • support GIS coordinates and layered, incident-specific maps

It also listed twelve desirable outcomes. They included resource details on demand, user queries and filters, real-time request status, open data formats, offline access to the last-known data, privacy, and pulling in data from every agency on scene.

What made it hard

  • No single radio fits every resource. A fire extinguisher or a first-aid kit needs a tag that’s small, cheap and runs for months on a battery, which rules out a cellular modem. A fire truck has power to spare and has to reach the cloud from anywhere. Responders already carry phones.
  • The newest network was the least proven. LTE-M, the low-power cellular standard built for devices like these, had only just launched in Canada. Modules were new and documentation was thin.
  • The radios had never been combined this way. Bluetooth Low Energy (BLE), LoRa and LTE-M were each mature on their own. Running all three together, with a gateway bridging between them, was the untested part.
  • The picture had to stay live. A command post needs positions that update in seconds, not a report after the fact. Data had to keep flowing even if part of the system failed.
  • The data is sensitive. Where responders and their equipment are during an incident isn’t something to put on a public cloud by default.

How we approached it

Match the radio to the resource. Instead of forcing one technology onto everything, we defined four device types. Each one matched what a resource can carry and how far its signal needs to travel.

Incident zoneLocal Tagon a kit or supply caseBLE + LoRaGatewayin a vehicleGPS + LTE-MDirect Tagon an asset or a personGPS + LTE-MPowered Taga responder's phoneGPS + LTEMomentrackprivate cloudLive incident mapBLE and LoRa,every secondLTE-M, batchedevery 5 secondsLTE-MLTEWebSocket

Research first, build second. We followed the V-model: design the concept, then prove each part against it. We ranked every requirement by how ready its technology already was. Anything we knew would work, such as assigning tags, printing labels or building reports, stayed out of scope. The six months went to the uncertain parts: radio interoperability, LTE-M on Canadian networks, and real-time visualization on a map.

Own the whole stack. The cloud ran on servers we operated ourselves, as stand-alone software with no dependency on AWS, Azure or Google Cloud. For the proof of concept we made one exception, Google Maps, and planned to replace it (see What came next).

Measure every hop. Every tag and gateway logged what it sent and received to an SD card. We compared those logs with what reached the database. That gave us a reliability figure for each link, not just for the system as a whole, and it showed exactly where data was being lost.


What we built

A data model for multi-agency incidents

Every agency on scene is a stakeholder. Each stakeholder owns resources of three kinds: assets (equipment, supplies, tools), people, and vehicles. A resource is tracked by a device attached to it or assigned to it. Devices are kept separate from resources so a tag can be put into inventory and later moved to a different resource. Each device carries sensors. Commanders draw incident zones on the map as polygons, and requests move resources into them.

owns

dispatches

for

tracked by

carries

1

1

1

*

0..*

*

Stakeholder

Resource

status

Device

Sensor

IncidentZone

Request

Asset

Person

Vehicle

LocalTag

DirectTag

PoweredTag

Gateway

Each resource type gets the devices that fit it:

  • Assets: a Local Tag or a Direct Tag
  • People: a Powered Tag or a Direct Tag
  • Vehicles: a Gateway, a Powered Tag or a Direct Tag

The devices

For a proof of concept, off-the-shelf hardware was the fastest way to test radios, not circuit boards. We built every device except the phone on Pycom’s FiPy, a single module with BLE, LoRa and LTE-M radios. Sensor shields went on top, and the firmware was written in MicroPython.

DeviceCarried byBuilt fromSensorsConnects over
Local Tagan asset: a kit, an extinguisher, a supply caseFiPy + Pysense shieldtemperature, humidity, air pressure, 3-axis accelerometer, light, batteryBLE and LoRa to a nearby gateway
Direct Tagan asset or a person, anywhere with coverageFiPy + Pytrack shieldGPS/GLONASS, 3-axis accelerometer, batteryLTE-M straight to the cloud
Powered Taga responderan Android phone running a sensor-collector appthe phone’s GPS and sensorsLTE straight to the cloud
Gatewaya vehicle, on vehicle powerFiPy + Pytrack shieldGPS/GLONASS, 3-axis accelerometerlistens for BLE and LoRa, relays over LTE-M

A Local Tag has no GPS by design. It stays small and cheap, and its position is the position of the gateway or phone that hears it. That’s less precise, but good enough to say which vehicle a defibrillator is in.

The radio layer

  • BLE: each Local Tag advertised its readings once a second in a custom packet modeled on Eddystone-TLM. The tag ran in peripheral-only mode to save power.
  • LoRa: raw LoRa rather than LoRaWAN, on 927 MHz in the US902–928 band used in Canada, at +20 dBm with 500 kHz bandwidth. Tags only transmitted, never listened, which saved power and kept the channel free.
  • LTE-M: we tested SIM cards from six providers: Telus, Rogers, Bell, Twilio, Hologram and iBasis. In early 2020, only Hologram, running on Rogers’ network, connected at all.
  • Payloads: a Local Tag’s full set of readings fit in 19 bytes over LoRa, and in a single BLE advertisement. Gateways and Direct Tags sent JSON to the cloud.

The private cloud

The cloud had seven services, each on its own virtual machine on a VMware server we ran ourselves. Each one could fail or be updated without taking the others down. If the API stopped, the IoT Hub kept accepting data from the field and the queue held it until the API was back.

Gatewaysand tagsIoT Hubauthenticate, throttle,validateProcessordecode payloads,type each readingData Access Layerbulk writesMongoDBreplica setAPIREST + SignalRWeb applive map, charts, reportsHTTP(S) or MQTT(S)RabbitMQRabbitMQRabbitMQ:data changedWebSocket
  • IoT Hub: the front door for every device. It took HTTP(S) and MQTT(S), checked each device’s credentials and security settings, and put the message on the queue.
  • Processor: decoded each payload into typed readings (sensor, value, unit and prefix).
  • Data Access Layer: the only service that wrote to the database. It used bulk writes to keep up with many small readings arriving at once.
  • Queue: RabbitMQ between every layer, with a topic for each kind of message.
  • Repository: MongoDB as a replica set of two data nodes and an arbiter, with geospatial queries for incident zones.
  • API: REST for queries and management, and SignalR to push live updates to browsers.
  • Web app: a single-page Vue application.

The live incident map

The web app is where it all came together. Users could:

  • define stakeholders, resources, devices and sensors
  • draw incident zones as polygons on the map
  • watch every resource move live, with standardized icons for trucks, ambulances, helicopters and field camps, map layers, and movement trails
  • click any resource to see its owner, status and latest sensor readings
  • browse each gateway’s and device’s raw data as it arrived, which is how we debugged the hardware
  • chart and filter sensor history, and export any report to CSV
  • create and track resource requests (basic management in the proof of concept)

To show a busy incident, we built a tag data simulator. It replayed an hour of real recorded device traffic, mixed with synthetic data, to the cloud every three seconds.

The request flow, as designed for the full system:

Resource ownerCommand postMomentrackResponder in the fieldRequest a resource for an incident zoneRequest waiting for approvalApproveDispatchAcknowledgeLive status and location of the resourceLive status and location of the resource

Proving it: ten tests

Every test compared what the devices logged locally with what arrived in the cloud.

#TestSetupResult
1Local Tag sensor read4 tags, one reading a second for 5 minutesa full sweep of all sensors in 541–550 ms
2BLE advertising4 tags, checked with a phone scannerpayload format validated on every tag
3BLE and LoRa at 1 m3 tags, every 2 seconds for 5 minutesLoRa: 535 of 536 received. BLE: 484 of 536
4LTE-M connection4 Direct Tags, 100 cycles a day for 5 days, SIMs from 6 providers7.5 s on average to attach and connect; only 1 provider of 6 worked
5LTE-M upload10 KB sent 100 times per tag17 kbps on average, about 4.6 s per 10 KB
6GPS fix4 stationary tags, one fix a secondcold start 117 s on average (45–244 s); later fixes 0.23 s
7Road test4 battery-powered Direct Tags in a moving vehicle, 5 runs of 30 minutes9,856 of 9,989 location reports reached the cloud (98.7%)
8Phone as Powered TagAndroid sensor-collector appcontinuous sensor stream received in the cloud
9Gateway, end to end3 Local Tags, 1 gateway at 3 m, 3 runs of 10 minutesLoRa: 95% of readings stored. BLE: 58%. Gateway to cloud: 188 of 188 batches (100%)
10Full integrationtags and gateways live, incident zones definedlive map with icons, layers, trails and sensor data: TRL 4

What the numbers told us

  • LoRa was the dependable short-range link. 95% of LoRa readings made it from tag to database.
  • BLE was the weak link, and per-hop logging showed where it failed. Only 3,846 of 6,480 BLE advertisements reached the gateway. Of those that did, 98% reached the database. Almost all the loss happened in the air between tag and gateway, so the fix belonged in the radio design, not the cloud.
  • LTE-M worked, but the market was young. 7.5 seconds to connect and 17 kbps are plenty for location updates. Only one SIM provider out of six connected, though.
  • GPS cold starts are slow. Two minutes on average to get a first fix meant a production tag would have to keep its GPS warm, or take fixes only when it moves.
  • One device misbehaved. One of the four Direct Tags was unstable in every road-test run. We sent its logs to the manufacturer to investigate.

My role

  • Led the bid. I led Momentaj’s proposal. DND’s evaluators gave it full marks for innovation, citing three significant improvements: multi-range radio for a reliable, highly available network; a dynamic dashboard; and battery management for low-power tags.
  • Designed the solution end to end. I owned the solution architecture across devices, radios and the cloud: the four device types, the choice of which radio to use where, and the layered, queue-based cloud.
  • Wrote code across the stack. I was hands-on in the device firmware, the IoT Hub, the Processor and Data Access Layer, and the API and web app.
  • Ran the project. I managed the team of nine through five milestones and reported to DND’s technical authority at Defence Research and Development Canada. The team included a hardware lead and radio designer, a senior .NET developer, a front-end developer, a UX/UI designer, a business analyst, and three junior hardware and firmware engineers, one of them a new graduate and one a University of Waterloo co-op student.
  • Validated the results. I analyzed the test logs and validated the results myself.
  • Led the Phase 2 proposal. I shaped the architecture and the plan for the Phase 2 proposal described below.

The cloud side reused the architecture I had already proven on uBeac, my IoT platform: ingest fast, queue everything, decode on the server, store, and stream live to the browser. That’s why a small team could have a seven-service private cloud running by the second milestone.


Outcome

  • TRL 3 → TRL 4 at both system and subsystem level
  • Five milestones out of five delivered on schedule, through the first COVID-19 lockdown. The final report went to DND on May 24, 2020.
  • A working end-to-end system: four device types, three radio technologies, seven cloud services and a live incident map
  • Measured baselines for the next phase: LoRa 95%, gateway to cloud 100%, LTE-M 7.5 s to connect at 17 kbps, 98.7% of road-test location reports delivered
  • Invited by Canada to propose Phase 2

What came next

In December 2020, Canada invited Momentaj to submit a Phase 2 proposal. Phase 2 offered up to two years and $1 million to turn the proof of concept into a working prototype that DND would own. We proposed an 18-month plan aimed at TRL 8:

  • Start with the responders. The plan began with five months of requirements work with the people who would use the system: interviews, workflow analysis and process models, all before any hardware design. This answered the main gap DND’s evaluators had found in our Phase 1 plan.
  • Real hardware. Custom tags and gateways, designed with Canadian electronics design firms. The plan called for three rounds of engineering prototypes, each tested against the specification and integrated with the cloud, plus enclosures and certification. Local Tags would run for months on coin cells, waking when they moved or on a timer.
  • Mobile apps. iOS and Android apps would let responders see nearby resources, send and acknowledge requests, and act as Powered Tags.
  • A cloud built for scale. Storage would be split in three: live data that expires automatically, an archive tuned for queries over time, and business records. Other planned pieces:
    • real-time streaming as its own service, with compression
    • encrypted device payloads
    • an integration agent to pull resource data from other agencies’ files, databases and systems
    • single sign-on with Active Directory and LDAP
    • English and French
    • drag-and-drop dashboards
  • Maps that run anywhere. Open-source mapping (Leaflet, OpenLayers and OpenStreetMap) hosted on-premises would replace Google Maps, so the system could work offline or on a closed network.
  • Two-way communication. Configuration and firmware updates would be sent over the air from the cloud to the devices.

The plan’s workstreams, with the time budgeted for each. They overlapped, so the whole plan ran 18 months:

  • Initiation: 0.5 month
  • Requirements with first responders: 5 months
  • Design: 3 months
  • Cloud application: 8 months
  • iOS and Android apps: 4 months
  • Hardware, in three engineering-prototype rounds: 10.5 months
  • Test, verification and integration: 4 months
  • Delivery, training and close: 2.5 months

DND didn’t go ahead with Phase 2 for this challenge, so Momentrack ended at the proof-of-concept stage.


What I learned

  • Measure every hop, not just the whole system. An end-to-end number would have said “BLE is unreliable.” Logging at each hop showed that nearly all the loss was in the air between tag and gateway, which told us exactly what to fix.
  • The newest link is the riskiest dependency. LTE-M was the right technology for the job, and in early 2020 one SIM provider out of six worked. When a design depends on a network that just launched, test the network before trusting the datasheet.
  • Listen to the evaluators, even after you’ve won. DND’s reviewers pointed out two gaps in our Phase 1 plan. It didn’t spend enough time with the first responders who would use the system, and it didn’t name radio interoperability as a risk. Our weakest test result came from the multi-radio gateway, the area they had flagged. Our Phase 2 plan started with the users.
  • Scope by uncertainty. Ranking every requirement by technology readiness let a team of nine prove the risky parts in six months, instead of spreading itself across features we already knew how to build.
  • Decide early what you won’t depend on. Running our own cloud was a deliberate choice for sensitive data. Google Maps was the one shortcut we took, and the Phase 2 plan had to budget time to undo it.

Technical deep dive

Payload formats

A Local Tag’s readings were packed into a few bytes so they fit in a single BLE advertisement or a short LoRa frame.

LoRa (19 bytes)

BytesFieldEncoding
0–5MAC addressdevice ID
6Humidity0–200, 0.5% per step
7Temperature, whole partsigned
8Temperature, fraction0–99, hundredths
9–10Air pressurevalue minus 50 kPa, most significant byte first
11–12Acceleration Xsigned 16-bit
13–14Acceleration Ysigned 16-bit
15–16Acceleration Zsigned 16-bit
17–18Batterymillivolts

BLE carried the same readings inside a standard advertisement: the BLE flags, a manufacturer-specific data block with a manufacturer ID, then a format byte and the sensor fields above. The tag’s MAC address came from the advertisement itself.

For the tests, both formats carried an extra 2-byte counter at the end. Comparing counters on the tag, the gateway and the database showed exactly which packets were lost, and where.

Gateways and Direct Tags sent JSON to the cloud. Each reading carried its sensor type, unit and unit prefix as codes, so the cloud could handle any sensor without a schema change:

[
  {
    "id": "1A:2A:3A:4A:5A:6A",
    "ts": 1542326605,
    "sensors": [
      { "id": "SignalStrength", "type": 17, "unit": 22, "prefix": 0,
        "data": { "rssi": -86, "txPower": -59 } },
      { "id": "Temperature", "type": 4, "unit": 2, "prefix": 0, "value": 26 }
    ]
  }
]

Firmware behavior

  • Local Tag: wake, read every sensor (about 0.55 s), advertise on BLE and LoRa, then deep-sleep for a configurable interval.
  • Direct Tag: attach to the LTE-M network, get a GPS fix, POST the reading as JSON, then idle.
  • Gateway: listen for BLE and LoRa on separate threads, collect everything heard into one JSON batch, and POST it to the cloud every 5 seconds.

The LTE-M connection lifecycle

Measured over 2,000 cycles (4 tags, 100 cycles a day, 5 days):

State changeAverage time
Attach to the network6.2 s
Open a data connection1.3 s
Disconnect6.2 s
Detach0.7 s

Every wake-up costs about 7.5 seconds of modem time before a single byte is sent. That is why reporting intervals and batching matter so much for battery life on LTE-M.

IoT Hub

The IoT Hub was built on ASP.NET Core 2.2, with MQTTnet handling MQTT. Every request passed through a chain of middleware:

  1. Identify the device. Tags are identified by MAC address, gateways by an ID set in their firmware. Validation data was cached in memory for 5 seconds, so the database wasn’t queried on every request.
  2. Throttle. Each device had a request limit, one a second by default.
  3. Apply the device’s security settings:
    • custom HTTP headers
    • an allow-list or block-list of IP addresses or ranges
    • an MQTT username and password
    • which protocols the device may use (HTTPS, MQTTS or both)
  4. Wrap the payload in a message envelope and put it on the queue.

Processor, data layer and queue

  • The Processor turned raw payloads into typed readings. Its decoders were C# functions compiled at runtime, so supporting a new payload format didn’t require a redeployment.
  • Services passed messages over RabbitMQ, with a topic for each type of object. Messages were serialized with MessagePack and compressed with LZ4.
  • The Data Access Layer used MongoDB bulk writes, with write concerns tuned for each scenario. On every change, it sent a message so the API could push the update to browsers.
  • MongoDB ran as a replica set of two data nodes and an arbiter. Its geospatial queries found the resources inside each incident zone.

API and web app

  • API: ASP.NET Core with REST endpoints, SignalR for live updates, JWT authentication through IdentityServer, and Swagger documentation.
  • Web app: Vue 2 with Vuex, Bootstrap-Vue, Highcharts and Google Maps. It used WebSocket for live data because browsers can’t speak MQTT or CoAP directly.
  • Map choice: we built test versions on both Google Maps and Leaflet before choosing Google Maps for the proof of concept.

Hosting and tooling

  • Every service ran on its own virtual machine on VMware ESXi 6.5, on our own server.
  • Firmware, API and UI source code, and test results, lived in Azure DevOps.
  • Every API was tested with Postman before any hardware was connected.

Stack

Pycom FiPy (BLE, LoRa, LTE-M) · Pysense and Pytrack shields · MicroPython · Android · C# · .NET Core / ASP.NET Core 2.2 · MQTTnet · RabbitMQ · MessagePack · MongoDB (replica set, geospatial queries) · IdentityServer · SignalR · Swagger · Serilog · Vue 2 · Vuex · Bootstrap-Vue · Highcharts · Google Maps · VMware ESXi · Azure DevOps

References