Skip to content

Brainish: Formalizing A Multimodal Language for Intelligence and Consciousness

Paul Pu Liang

arXiv Preprint Archive April 14, 2022 via arXiv

Summary

AI-generated from the abstract

A multimodal inner language called Brainish, combining words, images, audio, and sensations, is proposed as essential for machine consciousness. Building on the Conscious Turing Machine (CTM) model, the authors define Brainish's syntax and semantics and operationalize it through multimodal AI: unimodal encoders segment data, a coordinated representation space composes features for holistic meaning, and decoders produce predictions or raw data. A simple implementation of Brainish is evaluated on multimodal prediction and retrieval tasks using real-world image, text, and audio datasets. The authors argue that such an inner language is crucial for communication and coordination in achieving machine consciousness and intelligence.

Study at a glance

Characteristics Theoretical or philosophical paper Peer reviewed
Keywords Cs.ai Cs.cl Cs.lg
Key finding A multimodal inner language called Brainish, built on the Conscious Turing Machine model and operationalized through multimodal AI, is argued to be important for advances in machine models of intelligence and consciousness.

Abstract

Having a rich multimodal inner language is an important component of human intelligence that enables several necessary core cognitive functions such as multimodal prediction, translation, and generation. Building upon the Conscious Turing Machine (CTM), a machine model for consciousness proposed by Blum and Blum (2021), we describe the desiderata of a multimodal language called Brainish, comprising words, images, audio, and sensations combined in representations that the CTM's processors use to communicate with each other. We define the syntax and semantics of Brainish before operationalizing this language through the lens of multimodal artificial intelligence, a vibrant research area studying the computational tools necessary for processing and relating information from heterogeneous signals. Our general framework for learning Brainish involves designing (1) unimodal encoders to segment and represent unimodal data, (2) a coordinated representation space that relates and composes unimodal features to derive holistic meaning across multimodal inputs, and (3) decoders to map multimodal representations into predictions (for fusion) or raw data (for translation or generation). Through discussing how Brainish is crucial for communication and coordination in order to achieve consciousness in the CTM, and by implementing a simple version of Brainish and evaluating its capability of demonstrating intelligence on multimodal prediction and retrieval tasks on several real-world image, text, and audio datasets, we argue that such an inner language will be important for advances in machine models of intelligence and consciousness.

Comments

No comments yet.

Log in to comment