"Train Your AI with Agentic Vision" A next-generation ML training platform powered by Gemini 3 Flash. Built specifically for the Google Gemini 3 Hackathon.
Machine Learning development is often bottlenecked by the "Data Drudgery"—hours of manual labeling and expensive GPU compute. Agntic eliminates these barriers. By leveraging the multimodal reasoning of Gemini 3 Flash, we’ve created an autonomous pipeline that turns a simple video into a production-ready vision classifier in minutes, directly in the browser.
Agntic orchestrates a team of specialized AI agents working in a coordinated reasoning loop:
- 🎬 Video Analyzer (Gemini 3 Flash): Scans raw video sequences using Gemini's 1M token context window to discover unique archetypes and temporal context.
- 🔍 Vision Cropper (Agentic Grounding): Employs Agentic Vision with Code Execution to zoom into frames, draw high-precision bounding boxes, and extract training samples autonomously.
- 🧠 ML Engineer (TF.js & Logic): Audits dataset variety using Cosine Similarity matrices, suggests optimal hyperparameters, and monitors training health.
- 🚀 Deployer (Validation): Validates model performance, generates confusion matrices, and prepares the model for instant edge deployment.
| Landing Page | Studio Dashboard |
|---|---|
![]() |
![]() |
| Project Setup | Model Evaluation |
|---|---|
![]() |
![]() |
| Inference Mode | Model Gallery |
|---|---|
![]() |
![]() |
Unlike traditional object detectors, Agntic uses Gemini 3 Flash's spatial grounding to locate objects by time and coordinates (
Agntic performs Transfer Learning directly in the user's browser using TensorFlow.js with WebGL acceleration. This ensures total data privacy (no training data leaves the device) and allows for instant "Vibe-Check" verification.
The system implements Cosine Similarity analysis (via tf.matMul) to detect and remove redundant or near-duplicate frames. This automated curation ensures high dataset variety and prevents overfitting, making small datasets remarkably effective.
- Framework: Next.js 15+ (App Router, Standalone Output)
- AI Core: Google Gemini 3 Flash
- Neural Engine: TensorFlow.js (MobileNetV3 Feature Extractor)
- Cloud Infrastructure: Google Cloud Run
- Backend/DB: Firebase Firestore & Firebase Storage
- Media: FFmpeg (WASM) & Sharp
- Clone the repository:
git clone https://github.com/dzakwanalifi/agntic.git cd agntic - Install dependencies:
npm install
- Configure Environment Variables (
.env.local):GEMINI_API_KEY=your_key_here NEXT_PUBLIC_FIREBASE_API_KEY=... NEXT_PUBLIC_FIREBASE_PROJECT_ID=... NEXT_PUBLIC_FIREBASE_STORAGE_BUCKET=... # ... other firebase vars
- Run development server:
npm run dev
Deploy to Cloud Run:
gcloud builds submit --config cloudbuild.yaml .Built with ❤️ for the Google Gemini 3 Hackathon by Dzakwan Alifi.






