Skip to content

Latest commit

 

History

41 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

VibeSync: AI-Powered Mood & Activity DJ

VibeSync is an intelligent system that detects user activity and mood to play fitting music. It offers two distinct architectural approaches for comparison:

  1. VLM Version: Uses a unified Vision-Language Model (Qwen2.5-VL) to analyze frames and generate JSON outputs.
  2. YOLO+LLM Version: Uses YOLOv11 (Pose) + YOLOv12 (Objects) for detection, followed by an LLM (DeepSeek) for reasoning.

Two Versions

This project offers two implementations:

1. VLM Version (src/vlmversion/)

  • Uses Vision-Language Model (Qwen2.5-VL-32B) for unified scene analysis
  • New Features:
    • 🎭 Face Recognition with Multi-Pose Capture
    • 👤 Identity Persistence & YOLO Person Tracking
    • ⚡ Manual Trigger (Key 'T' for immediate scan)
    • 📊 Dual Progress Bars (Action Build-up + Stability Focus)
    • 🎨 Skeleton Visualization on Video
    • 🔄 Hysteresis Motion Detection (prevents flickering)
    • ⚙️ config.json for settings
  • Multi-Frame Temporal Analysis
  • Motion-based Smart Triggers
  • YouTube Music Streaming

2. YOLO+LLM Version (src/yolollmversion/)

  • Uses YOLO Pose + Object Detection + DeepSeek LLM
  • Lightweight approach without Vision Model
  • Good for group activities with multiple people
  • Genre-based music playback

Setup

📦 Prerequisites & Installation

  1. Install Python 3.10+

  2. Install CMAKE (Required for face recognition):

    pip install cmake
  3. Install dlib (if automatic installation fails):

    If pip install dlib fails during requirements installation, follow these steps:

    a. Make sure CMAKE is installed (see step 2 above)

    b. Download dlib from GitHub:

    c. Install dlib manually:

    cd path/to/dlib-master
    python setup.py install
  4. Install Remaining Dependencies:

    pip install -r requirements.txt
  5. Prepare Music Library: Create a Songs folder in the project root with subfolders for each activity. Add .mp3 or .wav files to them.

    VibeSync/
    ├── Songs/
    │   ├── Chilling/
    │   ├── Exercising/
    │   ├── Partying/
    │   └── Studying/
    

🔑 Configuration

Create a .env file in the project root (or src/vlmversion / src/yolollmversion depending on where you run from, but root is recommended if running via module path):

HUGGINGFACE_API_KEY=your_huggingface_key_here

(Note: The system uses Fireworks AI via OpenAI-compatible endpoints, ensure keys are set up for the respective providers if needed, or check specific llm_analyzer.py files for endpoint details.)

🚀 Running the System

Option 1: VLM Version (Unified)

This version sends visual frames to a multimodal model for holistic analysis.

python src/vlmversion/main.py
  • Controls: Press q to quit.
  • Config: src/vlmversion/user_profile.json (User preferences), src/vlmversion/activities.json (Activity list).

Option 2: YOLO + LLM Version (Modular)

This version uses object/pose detection followed by text-based reasoning.

python src/yolollmversion/main.py
  • Controls: Press q to quit.
  • Note: Requires yolo11n-pose.pt and yolo12n.pt (automatically downloaded by Ultralytics on first run).

📊 Benchmarking

Compare the accuracy and performance of both systems.

  1. Prepare Test Images: Place your ground truth images in: src/benchmarking/images/ (e.g., chilling1.jpg, studying2.jpg, etc.)

  2. Run Text Benchmark: Runs the systems on all images and saves a JSON report.

    python src/benchmarking/benchmark_system.py
  3. Run Visual Benchmark: Generates charts and confusion matrices comparing accuracy and speed.

    python src/benchmarking/benchmark_system_visual.py
    • Output: Plots are saved to src/benchmarking/benchmarking_plots/.

📂 Project Structure

  • src/vlmversion: Code for the Vision-Language Model approach.
  • src/yolollmversion: Code for the YOLO + LLM approach.
  • src/benchmarking: Evaluation scripts and datasets.
  • Songs: Local music library.

About

Activity Based Music Player

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages