This project implements a vision-based robot navigation system that allows a robot to recognize its current location and navigate toward a target using previously recorded images.
The system uses RootSIFT image descriptors, a K-Means visual vocabulary, VLAD feature aggregation, cosine similarity, a navigation graph, and shortest-path planning. :contentReference[oaicite:0]{index=0}
The program extracts SIFT features from each recorded image and converts them into RootSIFT descriptors. These local image features are then grouped using a K-Means visual vocabulary. :contentReference[oaicite:1]{index=1}
The descriptors are combined into one VLAD feature vector for each image. The VLAD vectors are normalized so that images can be compared using cosine similarity. :contentReference[oaicite:2]{index=2}
During preprocessing, the program creates a database containing the VLAD feature vector for every recorded navigation frame. SIFT descriptors and the K-Means codebook are cached to reduce loading time during later runs. :contentReference[oaicite:3]{index=3}
During navigation, the robot captures its current first-person view and compares it against the stored VLAD database.
The stored image with the highest cosine similarity is selected as the robot's estimated current location. :contentReference[oaicite:4]{index=4}
Each recorded image is represented as a node in a graph. Consecutive frames are connected with temporal edges based on the recorded movement action. :contentReference[oaicite:5]{index=5}
The target image is matched to the closest graph node, and NetworkX shortest-path planning is used to calculate a path from the robot's current location to the goal. :contentReference[oaicite:6]{index=6}
Autopilot follows the planned path by recovering the action associated with each graph edge.
The system supports:
- Forward movement
- Backward movement
- Left turns
- Right turns
- Automatic stopping when progress is no longer being made
Autopilot pauses after repeated localization or navigation failures to prevent the robot from continuing in the wrong direction. :contentReference[oaicite:7]{index=7}
The navigation panel displays:
- Live first-person camera view
- Best matching database image
- Target image
- Current node
- Goal node
- Number of remaining steps
- Recommended next movement
- Preview of upcoming path nodes
:contentReference[oaicite:8]{index=8}
| Key | Action |
|---|---|
| Arrow Up | Move forward |
| Arrow Down | Move backward |
| Arrow Left | Turn left |
| Arrow Right | Turn right |
| Space | Check in at the target |
| A | Enable or disable autopilot |
| Q | Display navigation information |
| Escape | Quit |
- Python
- OpenCV
- SIFT and RootSIFT
- VLAD image representation
- Scikit-learn K-Means
- NetworkX
- NumPy
- Pygame
Install the required Python packages:
pip install numpy opencv-python pygame networkx scikit-learn tqdm