Perceptron released Mk1.5, a 35B embodied-agent model with up to 4.7× faster request completion.
- First-person video is a very different problem from ordinary video understanding.
For a robot or a pair of smart glasses, hands constantly enter the frame, manipulate objects, disappear, return, and partially occlude what matters. The camera itself is moving too.
Perceptron trained Mk1.5 specifically on egocentric video, and very interestingly the model gets a 36.1-point jump on MMSearch when tools are enabled. Moves from 18.2 to 54.3 F1 when it can use web search, page reading, reverse-image search, and zoom.
This is a useful way to think about embodied intelligence too.
A robot does not need every possible capability encoded into its weights. It needs to recognize when its current information is insufficient, pick the right external capability, use it, then continue the task.