2 articles
Tencent's new Vision-Language Model introduces 4D reasoning that beats existing video understanding systems by over 20%, achieving 58.9% accuracy while maintaining strong general video comprehension.
A new AI method called Chain-of-Visual-Thought introduces continuous visual tokens to improve reasoning in vision-language models. Recent tests show accuracy gains of 3–16 percent across major benchmarks.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy