FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
A training-free method that selects a compact, representative set of visual tokens for long-video understanding. FLoC uses submodular optimization and works across video-language models without depending on the user query.
Preserving the details that matter in long video understanding
visual token compression ratio 1/32Green boxes: tokens selected by FLoC
What is the woman wearing during the summer sunset?
FLoC predictionA hat and sunglasses✓ Correct answerCompare methods on this example
| Method | Prediction | Result |
|---|---|---|
| FLoC | A hat and sunglasses | Correct |
| DivPrune | A dress and heels | Incorrect |
| TS-LLaVA | A dress and heels | Incorrect |
| Spectral Clustering | A dress and heels | Incorrect |
| K-means | A hat and sunglasses | Correct |
| K-medoids | A dress and heels | Incorrect |