Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎬 Agntic: Agentic Video-to-Classifier Studio

Next.js Gemini 3 Flash TensorFlow.js Cloud Run

Agntic Hero

"Train Your AI with Agentic Vision" A next-generation ML training platform powered by Gemini 3 Flash. Built specifically for the Google Gemini 3 Hackathon.

🔗 Quick Links


💡 The Inspiration

Machine Learning development is often bottlenecked by the "Data Drudgery"—hours of manual labeling and expensive GPU compute. Agntic eliminates these barriers. By leveraging the multimodal reasoning of Gemini 3 Flash, we’ve created an autonomous pipeline that turns a simple video into a production-ready vision classifier in minutes, directly in the browser.

🧠 The Agentic Workflow

Agntic orchestrates a team of specialized AI agents working in a coordinated reasoning loop:

  1. 🎬 Video Analyzer (Gemini 3 Flash): Scans raw video sequences using Gemini's 1M token context window to discover unique archetypes and temporal context.
  2. 🔍 Vision Cropper (Agentic Grounding): Employs Agentic Vision with Code Execution to zoom into frames, draw high-precision bounding boxes, and extract training samples autonomously.
  3. 🧠 ML Engineer (TF.js & Logic): Audits dataset variety using Cosine Similarity matrices, suggests optimal hyperparameters, and monitors training health.
  4. 🚀 Deployer (Validation): Validates model performance, generates confusion matrices, and prepares the model for instant edge deployment.

📸 Product Showcase

Landing Page Studio Dashboard
Landing Studio
Project Setup Model Evaluation
Project Eval
Inference Mode Model Gallery
Inference Models

🚀 Key Technical Highlights

1. Agentic Grounding with Code Execution

Unlike traditional object detectors, Agntic uses Gemini 3 Flash's spatial grounding to locate objects by time and coordinates ($[ymin, xmin, ymax, xmax]$). It then uses automated code execution to refine bounding boxes and extract normalized images, ensuring zero manual labeling.

2. Browser-Native Transfer Learning

Agntic performs Transfer Learning directly in the user's browser using TensorFlow.js with WebGL acceleration. This ensures total data privacy (no training data leaves the device) and allows for instant "Vibe-Check" verification.

3. SOTA Data Auditing

The system implements Cosine Similarity analysis (via tf.matMul) to detect and remove redundant or near-duplicate frames. This automated curation ensures high dataset variety and prevents overfitting, making small datasets remarkably effective.


🛠️ Tech Stack


🚥 Getting Started

Local Installation

  1. Clone the repository:
    git clone https://github.com/dzakwanalifi/agntic.git
    cd agntic
  2. Install dependencies:
    npm install
  3. Configure Environment Variables (.env.local):
    GEMINI_API_KEY=your_key_here
    NEXT_PUBLIC_FIREBASE_API_KEY=...
    NEXT_PUBLIC_FIREBASE_PROJECT_ID=...
    NEXT_PUBLIC_FIREBASE_STORAGE_BUCKET=...
    # ... other firebase vars
  4. Run development server:
    npm run dev

🚢 Deployment

Deploy to Cloud Run:

gcloud builds submit --config cloudbuild.yaml .

Built with ❤️ for the Google Gemini 3 Hackathon by Dzakwan Alifi.

About

Agentic Video-to-Classifier Studio. Next-generation ML training platform powered by Gemini 3 Flash and TensorFlow.js. Winner of the Google Gemini 3 Hackathon.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages