Hierarchical multi-layer preceptions (hMLPs) are a way to handle non-text data when training and inferencing in multimodal models.
Thinking Machines Lab used these instead of a vision encoder to project visual and audio data into their Inkling models.1