{"id":281844,"date":"2024-05-23T06:38:03","date_gmt":"2024-05-23T13:38:03","guid":{"rendered":"https:\/\/sftarticles.wpenginepowered.com\/es\/?p=332919"},"modified":"2025-07-01T16:27:09","modified_gmt":"2025-07-01T23:27:09","slug":"from-llama-to-chameleon-this-is-metas-new-multimodal-ai","status":"publish","type":"post","link":"https:\/\/cms-articles.softonic.io\/en\/from-llama-to-chameleon-this-is-metas-new-multimodal-ai\/","title":{"rendered":"From &#8220;llama&#8221; to &#8220;chameleon&#8221;: this is Meta&#8217;s new multimodal AI"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><strong>Meta<\/strong> has unveiled <strong>Chameleon<\/strong>, its new multimodal artificial intelligence designed to tackle the growing competition in the field of generative AI. Chameleon stands out for being <strong>natively multimodal<\/strong>, seamlessly integrating <strong>components from different modalities such as images, text, and code<\/strong>.<\/p>\n\n\n<div class=\"sc-card-program\">\r\n  <div class=\"sc-card-program__body\">\r\n    <div class=\"sc-card-program__row clearfix\">\r\n      <div class=\"sc-card-program__col-logo\">\r\n        <img decoding=\"async\" class=\"sc-card-program__img\" alt=\"Facebook\" src=\"https:\/\/images.sftcdn.net\/images\/t_app-icon-m\/p\/b2253bb6-9b53-11e6-8b9d-00163ed833e7\/219705778\/facebook-icon.png\" width=\"100px\" height=\"100px\">\r\n      <\/div>\r\n      <div class=\"sc-card-program__col-title\">\r\n        <span class=\"sc-card-program__title\">Facebook<\/span>\r\n        <a class=\"sc-card-program__button sc-card-program-internal\" href=\"https:\/\/facebook.en.softonic.com\/android\" target=\"_self\" rel=\"noopener noreferrer\">DOWNLOAD<\/a>\r\n      <\/div>\r\n      <div class=\"sc-card-program__col-rating\">\r\n        <svg class=\"rating-score__content\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" version=\"1.1\" x=\"0\" y=\"0\" viewbox=\"0 0 50 50\" enable-background=\"new 0 0 50 50\" xml:space=\"preserve\"><path class=\"rating-score__background rating-score--good\" fill=\"none\" stroke-width=\"6\" stroke-miterlimit=\"10\" d=\"M40 40c8.3-8.3 8.3-21.7 0-30s-21.7-8.3-30 0 -8.3 21.7 0 30\"><\/path><path class=\"rating-score__value rating-score__value--0\" fill=\"none\" stroke-width=\"6\" stroke-dashoffset=\"0\" stroke-miterlimit=\"10\" d=\"M40 40c8.3-8.3 8.3-21.7 0-30s-21.7-8.3-30 0 -8.3 21.7 0 30\"><\/path><text class=\"rating-score__number\" content=\"\" text-anchor=\"middle\" transform=\"matrix(1 0 0 1 25 31.0837)\" data-auto=\"app-user-score\"><\/text><\/svg>\r\n      <\/div>\r\n    <\/div>\r\n    <div class=\"sc-card-program__row\">\r\n      <span class=\"sc-card-program__description\"><\/span>\r\n    <\/div>\r\n    <div class=\"sc-card-program__row\">\r\n      <img decoding=\"async\" class=\"sc-card-program__bigpic\" src=\"\" onerror=\"this.style.display='none'\">\r\n    <\/div>\r\n    <a class=\"sc-card-program__link track-link sc-card-program-internal\" href=\"https:\/\/facebook.en.softonic.com\/android\" target=\"_self\" rel=\"noopener noreferrer\"><\/a>\r\n  <\/div>\r\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">According to the <a href=\"https:\/\/arxiv.org\/abs\/2405.09818v1\" target=\"_blank\" rel=\"noopener nofollow\" title=\"\">paper<\/a> published by the research team, Chameleon&#8217;s architecture allows for <strong>outstanding performance in tasks that require a deep understanding<\/strong> of both visual and textual information. Among Chameleon&#8217;s notable capabilities are image captioning and visual question answering (VQA), as well as its <strong>competitiveness in exclusively textual tasks<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Traditionally, multimodal models are created through a process known as <strong>&#8220;late fusion&#8221;<\/strong>, where the AI system processes different modalities separately and then merges the encodings for inference. However, <strong>this approach limits the ability of models to seamlessly integrate information across different modalities<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Chameleon has adopted an architecture of <strong>&#8220;early fusion based on mixed tokens&#8221;<\/strong>, which means that it has been designed from scratch to <strong>learn from an interleaved mixture of images, text, and other modalities<\/strong>. This methodology transforms images into discrete tokens, similar to how language models handle words, and uses a unified vocabulary of text, code, and image tokens.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Compared to similar models like <strong><a href=\"https:\/\/en.softonic.com\/articles\/this-is-how-ai-will-come-to-our-mobile-thanks-to-google-gemini-nano\" target=\"_blank\" rel=\"noopener\" title=\"\">Google Gemini<\/a><\/strong>, Chameleon offers <strong>a more cohesive integration of modalities during content generation<\/strong>, as it does not require specific components for each modality.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-large\"><img decoding=\"async\" src=\"https:\/\/articles-img.sftcdn.net\/sft\/articles\/auto-mapping-folder\/sites\/3\/2024\/05\/Meta-nueva-IA-1024x576-1-1024x576.jpg\" alt=\"\" class=\"wp-image-281848\" \/><\/figure>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\">The Chameleon training was carried out in two stages, using a vast dataset that includes <strong>4.4 trillion text tokens<\/strong>, image-text pairs, and interleaved text and image sequences. The Chameleon models, with <strong><a href=\"https:\/\/en.softonic.com\/articles\/the-new-ai-from-microsoft-is-very-very-big\" target=\"_blank\" rel=\"noopener\" title=\"\">7,000 and 34,000 billion parameters<\/a><\/strong>, were trained for over 5 million hours on 80 GB Nvidia A100 GPUs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The experiments showed that Chameleon can perform a wide range of text and multimodal tasks with market-leading performance. In VQA and image captioning tests, <strong>Chameleon-34B outperformed models like Flamingo, IDEFICS, and Llava-1.5<\/strong>. Additionally, <strong>it matched the performance of other models with fewer training examples<\/strong> in context and with smaller models.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Despite the complexity of multimodality, Chameleon remains competitive in text-only tasks, <strong>comparable to models like Mixtral 8x7B and Gemini-Pro in logical reasoning and reading comprehension tests<\/strong>. Researchers highlight that Chameleon unlocks new multimodal reasoning and generation capabilities, offering user-preferred results in documents that combine text and images in an interleaved manner.<\/p>\n\n\n<div class=\"sc-card-program\">\r\n  <div class=\"sc-card-program__body\">\r\n    <div class=\"sc-card-program__row clearfix\">\r\n      <div class=\"sc-card-program__col-logo\">\r\n        <img decoding=\"async\" class=\"sc-card-program__img\" alt=\"Facebook\" src=\"https:\/\/images.sftcdn.net\/images\/t_app-icon-m\/p\/b2253bb6-9b53-11e6-8b9d-00163ed833e7\/219705778\/facebook-icon.png\" width=\"100px\" height=\"100px\">\r\n      <\/div>\r\n      <div class=\"sc-card-program__col-title\">\r\n        <span class=\"sc-card-program__title\">Facebook<\/span>\r\n        <a class=\"sc-card-program__button sc-card-program-internal\" href=\"https:\/\/facebook.en.softonic.com\/android\" target=\"_self\" rel=\"noopener noreferrer\">DOWNLOAD<\/a>\r\n      <\/div>\r\n      <div class=\"sc-card-program__col-rating\">\r\n        <svg class=\"rating-score__content\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" version=\"1.1\" x=\"0\" y=\"0\" viewbox=\"0 0 50 50\" enable-background=\"new 0 0 50 50\" xml:space=\"preserve\"><path class=\"rating-score__background rating-score--good\" fill=\"none\" stroke-width=\"6\" stroke-miterlimit=\"10\" d=\"M40 40c8.3-8.3 8.3-21.7 0-30s-21.7-8.3-30 0 -8.3 21.7 0 30\"><\/path><path class=\"rating-score__value rating-score__value--0\" fill=\"none\" stroke-width=\"6\" stroke-dashoffset=\"0\" stroke-miterlimit=\"10\" d=\"M40 40c8.3-8.3 8.3-21.7 0-30s-21.7-8.3-30 0 -8.3 21.7 0 30\"><\/path><text class=\"rating-score__number\" content=\"\" text-anchor=\"middle\" transform=\"matrix(1 0 0 1 25 31.0837)\" data-auto=\"app-user-score\"><\/text><\/svg>\r\n      <\/div>\r\n    <\/div>\r\n    <div class=\"sc-card-program__row\">\r\n      <span class=\"sc-card-program__description\"><\/span>\r\n    <\/div>\r\n    <div class=\"sc-card-program__row\">\r\n      <img decoding=\"async\" class=\"sc-card-program__bigpic\" src=\"\" onerror=\"this.style.display='none'\">\r\n    <\/div>\r\n    <a class=\"sc-card-program__link track-link sc-card-program-internal\" href=\"https:\/\/facebook.en.softonic.com\/android\" target=\"_self\" rel=\"noopener noreferrer\"><\/a>\r\n  <\/div>\r\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Meta has unveiled Chameleon, its new multimodal artificial intelligence designed to tackle the growing competition in the field of generative AI. Chameleon stands out for being natively multimodal, seamlessly integrating components from different modalities such as images, text, and code. According to the paper published by the research team, Chameleon&#8217;s architecture allows for outstanding performance &hellip; <a href=\"https:\/\/cms-articles.softonic.io\/en\/from-llama-to-chameleon-this-is-metas-new-multimodal-ai\/\" class=\"more-link\">Continue reading<span class=\"screen-reader-text\"> &#8220;From &#8220;llama&#8221; to &#8220;chameleon&#8221;: this is Meta&#8217;s new multimodal AI&#8221;<\/span><\/a><\/p>\n","protected":false},"author":9256,"featured_media":281846,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","wpcf-pageviews":1},"categories":[1015],"tags":[2353],"usertag":[],"vertical":[],"content-category":[],"class_list":["post-281844","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-news","tag-app-subdomain-redirectionfacebook"],"aioseo_notices":[],"_links":{"self":[{"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/posts\/281844","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/users\/9256"}],"replies":[{"embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/comments?post=281844"}],"version-history":[{"count":1,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/posts\/281844\/revisions"}],"predecessor-version":[{"id":313025,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/posts\/281844\/revisions\/313025"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/media\/281846"}],"wp:attachment":[{"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/media?parent=281844"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/categories?post=281844"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/tags?post=281844"},{"taxonomy":"usertag","embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/usertag?post=281844"},{"taxonomy":"vertical","embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/vertical?post=281844"},{"taxonomy":"content-category","embeddable":true,"href":"https:\/\/cms-articles.softonic.io\/en\/wp-json\/wp\/v2\/content-category?post=281844"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}