VibeSync is an intelligent system that detects user activity and mood to play fitting music. It offers two distinct architectural approaches for comparison:
- VLM Version: Uses a unified Vision-Language Model (Qwen2.5-VL) to analyze frames and generate JSON outputs.
- YOLO+LLM Version: Uses YOLOv11 (Pose) + YOLOv12 (Objects) for detection, followed by an LLM (DeepSeek) for reasoning.
This project offers two implementations:
- Uses Vision-Language Model (Qwen2.5-VL-32B) for unified scene analysis
- New Features:
- 🎭 Face Recognition with Multi-Pose Capture
- 👤 Identity Persistence & YOLO Person Tracking
- ⚡ Manual Trigger (Key 'T' for immediate scan)
- 📊 Dual Progress Bars (Action Build-up + Stability Focus)
- 🎨 Skeleton Visualization on Video
- 🔄 Hysteresis Motion Detection (prevents flickering)
- ⚙️ config.json for settings
- Multi-Frame Temporal Analysis
- Motion-based Smart Triggers
- YouTube Music Streaming
- Uses YOLO Pose + Object Detection + DeepSeek LLM
- Lightweight approach without Vision Model
- Good for group activities with multiple people
- Genre-based music playback
-
Install Python 3.10+
-
Install CMAKE (Required for face recognition):
pip install cmake
-
Install dlib (if automatic installation fails):
If
pip install dlibfails during requirements installation, follow these steps:a. Make sure CMAKE is installed (see step 2 above)
b. Download dlib from GitHub:
- Go to https://github.com/davisking/dlib
- Click "Code" → "Download ZIP"
- Extract the zip file (should be named
dlib-master)
c. Install dlib manually:
cd path/to/dlib-master python setup.py install -
Install Remaining Dependencies:
pip install -r requirements.txt
-
Prepare Music Library: Create a
Songsfolder in the project root with subfolders for each activity. Add.mp3or.wavfiles to them.VibeSync/ ├── Songs/ │ ├── Chilling/ │ ├── Exercising/ │ ├── Partying/ │ └── Studying/
Create a .env file in the project root (or src/vlmversion / src/yolollmversion depending on where you run from, but root is recommended if running via module path):
HUGGINGFACE_API_KEY=your_huggingface_key_here(Note: The system uses Fireworks AI via OpenAI-compatible endpoints, ensure keys are set up for the respective providers if needed, or check specific llm_analyzer.py files for endpoint details.)
This version sends visual frames to a multimodal model for holistic analysis.
python src/vlmversion/main.py- Controls: Press
qto quit. - Config:
src/vlmversion/user_profile.json(User preferences),src/vlmversion/activities.json(Activity list).
This version uses object/pose detection followed by text-based reasoning.
python src/yolollmversion/main.py- Controls: Press
qto quit. - Note: Requires
yolo11n-pose.ptandyolo12n.pt(automatically downloaded by Ultralytics on first run).
Compare the accuracy and performance of both systems.
-
Prepare Test Images: Place your ground truth images in:
src/benchmarking/images/(e.g.,chilling1.jpg,studying2.jpg, etc.) -
Run Text Benchmark: Runs the systems on all images and saves a JSON report.
python src/benchmarking/benchmark_system.py
-
Run Visual Benchmark: Generates charts and confusion matrices comparing accuracy and speed.
python src/benchmarking/benchmark_system_visual.py
- Output: Plots are saved to
src/benchmarking/benchmarking_plots/.
- Output: Plots are saved to
src/vlmversion: Code for the Vision-Language Model approach.src/yolollmversion: Code for the YOLO + LLM approach.src/benchmarking: Evaluation scripts and datasets.Songs: Local music library.