DAVID NKWO

Robotics & Physical AI Engineer

David Nkwo

Building intelligent robotic systems that operate reliably in real-world environments.

My work spans autonomy, perception, manipulation, robot learning, ROS 2, multimodal AI and physical systems integration, with a systems-level focus from perception and planning through safety, recovery and deployment.

Autonomous Systems · Robot Learning · Manipulation · Systems Architecture · Integration · Technical Leadership

Selected work

Selected Work

Selected robotics and Physical AI work spanning physical hardware integration, autonomous systems, robot perception, manipulation and multimodal intelligence.

01In Progress · System Definition & Pre-Build

Physical AI Autonomous Delivery Robot

A 28-week Physical AI creation and engineering project focused on building a physical, language-conditioned autonomous delivery platform.

The system is being designed as one cumulative robotics platform combining ROS 2 hardware integration, grounded perception, manipulation, autonomous navigation, robot learning, semantic scene understanding, safety, recovery and reproducible system validation.

Current phase

System architecture, requirements, execution planning and hardware selection are complete. Physical implementation is the next stage.

Planned system architecture
Planned system architecture — design schematic, not a build photo.

Mobile platform

Hiwonder MentorPi A1 Advanced

Raspberry Pi 5

Manipulator

Hiwonder NexArm

Physical bring-up platform

Adeept 5-DOF Robotic Arm

Roadmap

  1. 01

    System Definition

    ● Complete

  2. 02

    Physical Foundation

    ○ Next

  3. 03

    Grounding → Action MVP

  4. 04

    Robot Learning

  5. 05

    Scene Reasoning

  6. 06

    Human-Aware Autonomy

  7. 07

    Integration & Validation

The project builds on earlier work in autonomous delivery robotics, adaptive visual grounding and physical ROS 2 hardware integration.

Instruction“the person near the elevator”
Person / lobby grounding result
Grounding output — person in a lobby scene.
02Deployed · Robot Perception

Adaptive Visual Grounding for Robot Vision

A deployed robot-vision system that converts natural-language instructions into localized visual targets across still images, recorded video and browser-camera inputs.

The system parses objects, attributes, quantities and spatial relationships, then routes requests across YOLO, OWL-ViT and GPT-guided OWL-ViT while producing structured bounding boxes, confidence, timing information and annotated outputs.

After comparing multiple grounding approaches, I designed and implemented a custom GPT Vision-guided OWL-ViT pipeline to improve localization accuracy while supporting instruction-aware selection, refinement and tracking.

  • YOLO
  • OWL-ViT
  • GPT
  • Visual Grounding
Table / flower relationship grounding
Relationship grounding — table / flower
Door / sign object and attribute grounding
Object & attribute grounding — door / sign
Second-closest package spatial / ordinal grounding
Spatial / ordinal grounding — second-closest package
Grey-door grounding and sustained tracking demonstration
Robot target grounding and tracking
Correct-package-on-chair grounding
Hidden-package spatial grounding
105
Evaluation images
0.473
Mean IoU
78.1%
Success @ IoU ≥ 0.25
57.1%
Success @ IoU ≥ 0.50

Results for the strongest configuration: the custom GPT Vision-guided OWL-ViT pipeline. The custom pipeline achieved the strongest mean IoU at 0.473, compared with OWL-ViT at 0.344, GPT Vision at 0.300 and Grounding DINO at 0.178 on the shared 105-image evaluation set.

03Physical Robotics · Completed · Demo Video Pending

ROS 2 Hardware Integration & Robotic Arm Bring-Up

Physical demo video coming soon (placeholder graphic)Physical demo video coming soon

Control path

  1. ROS 2
  2. C++ Hardware Interface / ros2_control
  3. joint_trajectory_controller
  4. Serial Protocol
  5. MCU Firmware
  6. Servo Control
  7. Physical Motion

Built a complete ROS 2-to-hardware control path for a physical robotic arm, integrating higher-level robot software with embedded firmware and physical actuation.

The system connects ROS 2 trajectory commands through a C++ hardware layer, serial protocol and MCU firmware to bounded servo control, while exposing robot state, TF, diagnostics and safety behaviour back into the ROS 2 environment.

The integration includes bounded joint commands, malformed-command rejection, communication timeouts, watchdog behaviour, diagnostics, calibration management and defined behaviour on communication loss.

Open-loop state

The arm does not provide encoder-based joint-position feedback. ROS joint state therefore represents the last accepted commanded estimate rather than verified physical proprioception.

RViz capture pending (placeholder graphic)Capture pending
RViz — capture pending
MoveIt capture pending (placeholder graphic)Capture pending
MoveIt — capture pending
TF capture pending (placeholder graphic)Capture pending
TF — capture pending
Diagnostics capture pending (placeholder graphic)Capture pending
Diagnostics — capture pending
  • ROS 2
  • C++
  • ros2_control
  • joint_trajectory_controller
  • Serial
  • MCU Firmware
  • URDF / Xacro
  • TF
  • robot_state_publisher
  • /joint_states
  • MoveIt 2
  • Watchdogs
  • Diagnostics
04ROS 2 · Autonomy · Perception · Manipulation

Autonomous Delivery Robotics

Multi-phase autonomous-delivery robotics work integrating ROS 2 navigation, robot perception, localization, manipulation and environment interaction.

The completed work spans single-floor autonomous navigation, package identification and localization, robotic pick-and-place, elevator perception and interaction, semantic indoor localization and language-conditioned visual grounding.

An integrated robotics programme demonstrated through validated modules — not a single production system.

M01

Autonomous Navigation

Nav2 and AMCL-based room-to-room navigation within a structured residential environment.

M02

Package Perception

OCR-based label reading, YOLO-based package awareness and ArUco pose estimation for manipulation.

M03

Manipulation

OpenManipulator and MoveIt 2 integration for simulated package interaction.

M04

Elevator Perception

YOLO-based elevator door-state detection combined with temporal decision logic for safer elevator-entry triggering.

M05

Button Interaction

Elevator button detection and classification connected to inverse-kinematics-based manipulator targeting.

M06

Semantic Localization

ResNet18-based classification of indoor locations using real building footage.

Simulated address-label recognition publishing the detected label “Apt 48”

M07

OCR Label Reading

Simulated address-label recognition publishing the detected label “Apt 48”.

Custom 2D Gymnasium/Pygame simulation — blue: robot, green: goal, black squares: moving obstacles, red: 24 LiDAR rays

M08

DQN Dynamic-Obstacle Navigation

A DQN policy uses 24 forward LiDAR rays and a relative-goal vector to reach the target while avoiding moving obstacles in a custom Gymnasium/Pygame simulation. 82% success across 100 zero-exploration evaluation episodes in the 20-moving-obstacle stress test.

100%
Elevator entry trigger success

Tested simulation conditions

0
False-positive entry triggers

Tested simulation conditions

0.778
Semantic localization test accuracy
05Multimodal AI · Data Systems

Vision-Language Model Evaluation Pipeline

Built a reproducible evaluation pipeline for benchmarking vision-language models across standardized multimodal question-answering datasets.

The system uses PySpark and Parquet to normalize benchmark records, provides model interfaces for BLIP-2, InstructBLIP and LLaVA-style workflows, performs answer-ranking evaluation and produces comparative results and heatmaps.

  • PySpark
  • Parquet
  • BLIP-2
  • InstructBLIP
  • LLaVA
Evaluation heatmap
Evaluation heatmap from actual results.
Evaluation pipeline diagram
Evaluation pipeline
Sample evaluation record
Actual sample evaluation record
Model comparison table
Model comparison from supplied results
10,301
ScienceQA records
1,000
VMCBench records
11,301
Evaluation items per model
49.11%
InstructBLIP weighted accuracy
45.16%
BLIP-2 weighted accuracy
Supporting work

Additional Technical Work

Supporting work across computer vision, reinforcement learning, cloud systems and applied AI.

Computer Vision · Temporal ML

ASL Video Sign Recognition

End-to-end video-recognition prototype using WLASL data, MediaPipe hand landmarks, temporal PyTorch models and a deployed Streamlit interface.

The application supports curated clips and exploratory uploaded videos, returning top-k predictions, confidence values and extracted landmark information.

Visual Grounding · Benchmarking

CLIP vs Grounding DINO

Referring-expression grounding experiment comparing a lightweight CLIP-based regression model against Grounding DINO across 1,150 test pairs.

Evaluation included Acc@0.5, mIoU, Acc@0.75, center error, inference time and category-level analysis.

RL · Control

Reinforcement Learning Systems

Implementations of dynamic programming, Monte Carlo control, SARSA, Q-learning, Expected SARSA, Tree Backup and REINFORCE across FrozenLake, CartPole and MountainCar.

Data · Cloud · ML

Cloud & Distributed ML Systems

Distributed-data and machine-learning work spanning Hadoop, PySpark, Spark MLlib, Azure Data Factory, Stream Analytics, IoT pipelines and Azure ML.

LLMs · RAG · Analytics

AI Strategy & Multi-Agent RAG

Applied AI strategy project combining market and policy analysis with a multi-agent retrieval-augmented generation system using OpenAI embeddings, FAISS, LangChain and CrewAI.

Systems engineering

Systems Engineering in Live Environments

My systems background extends beyond robotics into live technical environments where hardware, software, networking, signal flow and people all have to work together reliably under time constraints. This experience has included technical scoping, procurement, system architecture, installation, integration, configuration, troubleshooting and live operation.

Job 01 · Completed

System Design · Procurement · Installation · Commissioning

Church AV, Broadcast & Livestream Integration

Scoped, procured, installed and integrated a complete church audio, video, display and livestream system for weekly services while the wider building infrastructure was still being completed.

The project required coordinating installation around changing site availability and other active construction work, sequencing each subsystem so work could continue without blocking other trades while still meeting a constrained delivery schedule.

Commissioned-system reel

Motion montage assembled from project stills — not continuous service footage.

01 · Before

Before installation

02 · Build / installation

Installation and infrastructure in progress

03 · Commissioned / in use

Completed system in the service environment

04 · Broadcast / livestream

Live production and broadcast control

Visual systems

5 sanctuary TVs

2 main · 2 stage foldback · 1 announcements

1 LED screen

3 children's-area TV connections

Audio system

2 L/R main PA · 2 centre fills

2 stage monitors

Choir + backup choir mics

Guitar · keyboard · drums

In-ear / choir monitoring

Broadcast

2 PTZ cameras

OBS livestreaming

YouTube delivery

Operational work

Service audio EQ and mixing, display routing, camera configuration, livestream integration and live-service operation.

The project required more than equipment installation: it involved requirements definition, procurement decisions, physical installation, system integration, sequencing around other infrastructure work, commissioning and operational validation under deadline pressure.

Job 02 · Completed

Church AV, Broadcast & Livestream Integration

Completed AV, production and livestream integration comprising two primary display TVs plus one rear display TV, three PTZ cameras with a dedicated controller, an OBS-based livestream workflow, two main PA speakers, two stage monitors, choir audio support and lighting integration.

Displays
2 primary TVs + 1 rear/back display TV
Video capture & control
3 PTZ cameras + dedicated controller
Production & streaming
OBS + livestream integration
Main audio
2 main PA speakers
Stage monitoring
2 stage monitors
Additional integration
Choir audio support + lighting

Completed system

Completed AV integration — after
Encore · Supervision & show delivery

Live Event Technical Leadership

Additional live-event technical leadership experience spanning show preparation, room readiness, signal-flow verification, equipment coordination, operator alignment, troubleshooting and end-to-end delivery in time-sensitive environments.

Live event show — technical supervision
Live event show — show delivery
Configured conference presentation room with dual projection screens
Conference room and AV layout prepared for an event

The work reinforces the same systems habits applied in robotics: define interfaces, verify the full chain, coordinate people and technology, anticipate failure points and maintain reliable operation under real constraints.

Engineering approach

From Prototype to Reliable System

I approach robotics as an end-to-end product and systems problem rather than a collection of isolated algorithms. That means defining interfaces clearly, integrating software with physical hardware, instrumenting system behaviour, testing failure modes, maintaining bounded recovery and fallback behaviour, and expanding autonomy only when measured performance supports it.

  1. 01

    Define

  2. 02

    Architect

  3. 03

    Build

  4. 04

    Integrate

  5. 05

    Instrument

  6. 06

    Validate

  7. 07

    Deploy

  8. 08

    Improve

Measured capability

A command being accepted is not the same as the physical outcome being verified.

Bounded autonomy

Intelligent behaviour should operate inside explicit capability, safety, recovery and fallback boundaries.

Progressive complexity

Increase task and environmental complexity only after the underlying capability is sufficiently reliable.

Current

Independent Robotics Systems Consultant

Robotics systems, product definition and prototype strategy

I support early-stage robotics and automation initiatives by turning operational problems into executable technical plans. My work spans feasibility analysis, requirements definition, system architecture, platform selection, prototype planning and implementation strategy.

I translate real workflows into hardware and software requirements, subsystem boundaries, sensing and compute needs, cost and risk trade-offs, build-versus-buy decisions, and phased development roadmaps. I also define integration and prototype-validation strategies that surface technical unknowns early and establish measurable gates for further investment.

Focus areas

  • Feasibility and requirements definition
  • Robotics system and subsystem architecture
  • Hardware, sensing and compute selection
  • Build-versus-buy and cost/risk analysis
  • Prototype and integration planning
  • Validation criteria and phased product roadmaps
About

About

Portrait of David Nkwo

I am a Robotics and Physical AI Engineer focused on building intelligent robotic systems that operate reliably in real-world environments.

My work spans autonomy, perception, manipulation, robot learning, ROS 2, multimodal AI and VLA systems, with a systems-level focus from perception and planning through safety, recovery and deployment.

I bring several years of technical leadership and systems integration experience across architecture, deployment, troubleshooting, reliability and team coordination. This background positions me to lead technical direction across multidisciplinary robotics systems, from architecture and technology decisions through integration, validation and deployment.

I approach robotics as an end-to-end product and systems problem, with particular focus on Physical AI, manipulation, autonomous systems and service robotics.

Education

University of Toronto

Master of Engineering

Robotics, AI & Data Analytics

University of Ottawa

Bachelor of Applied Science

Mechanical Engineering

Let's connect

Interested in robotics, Physical AI, autonomous systems, technical collaboration or building intelligent physical products?