<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Google on IT Space</title><link>https://blog.jiatool.com/en/tags/Google/</link><description>Recent content in Google on IT Space</description><generator>Hugo -- gohugo.io</generator><language>en</language><managingEditor>jia@jiatool.com (Jia)</managingEditor><webMaster>jia@jiatool.com (Jia)</webMaster><copyright>&amp;copy;{year}, Jia All Rights Reserved</copyright><lastBuildDate>Fri, 13 Dec 2024 22:10:00 +0800</lastBuildDate><atom:link href="https://blog.jiatool.com/en/tags/Google/index.xml" rel="self" type="application/rss+xml"/><item><title>Gemini 2.0 Flash is Free to play, Real-Time Video Chat, Image Understanding</title><link>https://blog.jiatool.com/en/posts/gemini_2_flash/</link><pubDate>Fri, 13 Dec 2024 22:10:00 +0800</pubDate><author>jia@jiatool.com (Jia)</author><atom:modified>Sat, 14 Dec 2024 14:30:00 +0800</atom:modified><guid>https://blog.jiatool.com/en/posts/gemini_2_flash/</guid><description>(This article was translated by AI and then reviewed by a human.) Preface Two days ago (12/11), Google announced Gemini 2.0, saying it is a new model specially created for the agentic era✨, and it was the first to launch the &amp;ldquo;Gemini 2.0 Flash (Experimental)&amp;rdquo; model. To showcase the new capabilities of Gemini 2.0, Google has an interface in Google AI Studio for everyone to try</description><content:encoded>&lt;p>&lt;em>(This article was translated by AI and then reviewed by a human.)&lt;/em>&lt;/p>
&lt;h2 id="preface">Preface&lt;/h2>
&lt;p>Two days ago (12/11), Google announced Gemini 2.0, saying it is a new model specially created for the agentic era✨, and it was the first to launch the &amp;ldquo;Gemini 2.0 Flash (Experimental)&amp;rdquo; model.&lt;/p>
&lt;br/>
&lt;p>To showcase the new capabilities of Gemini 2.0, Google has an interface in &lt;a href="https://aistudio.google.com/" target="_blank" rel="noopener">
Google AI Studio
&lt;/a> for everyone to try out.&lt;/p>
&lt;p>In addition to the powerful functions of real-time video + voice interaction similar to the previous demo's Project Astra, the new model includes spatial image understanding, video analysis, and map integration, among other exciting features.&lt;br />
Let's dive in for a quick experience!&lt;/p>
&lt;br/>
&lt;p>By the way,&lt;br />
* Just a day later, OpenAI rolled out its advanced voice mode with similar real-time video features (including screen sharing). The competition is heating up! XD&lt;br />
* Google also introduced a &lt;a href="https://blog.google/products/gemini/google-gemini-deep-research/" target="_blank" rel="noopener">
Deep Research tool
&lt;/a>, which can search the web on your behalf (and may collect more than a 100 sources), analyze it, and compile it into research reports.&lt;/p>
&lt;br/>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini2.0.jpg" alt="Gemini 2.0" data-caption="Gemini 2.0" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='650px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:650px;height:;"/>
&lt;figcaption style="text-align: center;">
Gemini 2.0
&lt;/figcaption>
&lt;/figure>
&lt;br/>
&lt;!--adsense-->
&lt;br/>
&lt;h2 id="gemini-20">Gemini 2.0&lt;/h2>
&lt;p>Gemini 2.0 Flash has outperformed 1.5 Pro in key benchmark tests, with speeds twice as fast as 1.5 Pro. It supports inputs like text, images, videos, and audio, and now even outputs native images and audio!&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_benchmarks.jpg" alt="Comparison of Gemini 2.0 Flash, 1.5 Pro, and 1.5 Flash" data-caption="Comparison of Gemini 2.0 Flash, 1.5 Pro, and 1.5 Flash" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='650px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:650px;height:;"/>
&lt;figcaption style="text-align: center;">
Comparison of Gemini 2.0 Flash, 1.5 Pro, and 1.5 Flash
&lt;/figcaption>
&lt;/figure>
&lt;p>Gemini 2.0 Flash is now available on Google AI Studio and Vertex AI. Global Gemini users can also access 2.0 Flash via desktop or mobile web platforms, with the feature set to launch on the Gemini mobile app soon ~&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_select.jpg" alt="The Gemini chat interface allows selecting models" data-caption="The Gemini chat interface allows selecting models" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='450px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:450px;height:;"/>
&lt;figcaption style="text-align: center;">
The Gemini chat interface allows selecting models
&lt;/figcaption>
&lt;/figure>
&lt;ul>
&lt;li>&lt;a href="https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/" target="_blank" rel="noopener">
Google Official Introduction to Gemini 2.0
&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://developers.googleblog.com/en/the-next-chapter-of-the-gemini-era-for-developers/" target="_blank" rel="noopener">
Google Official Overview of New Features in Gemini 2.0 Flash
&lt;/a>&lt;/li>
&lt;/ul>
&lt;br/>
&lt;h2 id="google-ai-studio">Google AI Studio&lt;/h2>
&lt;h3 id="stream-realtime">Stream Realtime&lt;/h3>
&lt;p>&lt;a href="https://aistudio.google.com/live" target="_blank" rel="noopener">
Stream Realtime | Google AI Studio
&lt;/a>&lt;/p>
&lt;p>Here, you can use a &amp;quot;camera + microphone&amp;quot; to interact with Gemini in real-time, similar to the Project Astra demo shown earlier (just not as advanced yet 😅).&lt;/p>
&lt;p>For example, you can walk around with your phone and ask, &amp;quot;What is this?&amp;quot; or &amp;quot;What is that?&amp;quot; You can also share your screen to ask for help, like how to operate something. The range of applications is very broad.&lt;br />
Since Gemini 2.0 is natively multimodal, it supports both voice input and output. It can adapt its tone and speed based on the context, and you can even interrupt it mid-response.&lt;/p>
&lt;p>It's hard to fully show this through text, so feel free to try it out yourself!&lt;/p>
&lt;p>* Unfortunately, it currently doesn't support Chinese voice output. However, you can still ask it questions in Chinese and have it reply in English (or set it to reply in text, so it can respond in Chinese).&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_stream_realtime.jpg" alt="Gemini 2.0 Stream Realtime" data-caption="Gemini 2.0 Stream Realtime" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Gemini 2.0 Stream Realtime
&lt;/figcaption>
&lt;/figure>
&lt;iframe width="672" height="378" src="https://www.youtube.com/embed/9hE5-98ZeCg?si=7Q4qOkwc4A8VS5DI" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;br/>
&lt;h3 id="starter-apps">Starter Apps&lt;/h3>
&lt;p>&lt;a href="https://aistudio.google.com/starter-apps" target="_blank" rel="noopener">
Starter Apps | Google AI Studio
&lt;/a>&lt;/p>
&lt;p>On this page, Google has created three apps to let you experience Gemini 2.0’s capabilities in understanding images, analyzing videos, and integrating maps.&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_starter_apps.jpg" alt="Gemini 2.0 Starter Apps" data-caption="Gemini 2.0 Starter Apps" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Gemini 2.0 Starter Apps
&lt;/figcaption>
&lt;/figure>
&lt;p>* If you are a program developer and want to manually write code for interaction, you can refer to the official &lt;a href="https://github.com/google-gemini/starter-applets" target="_blank" rel="noopener">
GitHub project example
&lt;/a>&lt;/p>
&lt;br/>
&lt;h4 id="spatial-understanding">Spatial Understanding&lt;/h4>
&lt;p>Spatial Understanding can identify objects in an image. For example, you can ask it to find a magic wand in a picture:&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/spatial_understanding_01.jpg" alt="Find a magic wand - Spatial Understanding" data-caption="Find a magic wand - Spatial Understanding" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='800px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:800px;height:;"/>
&lt;figcaption style="text-align: center;">
Find a magic wand - Spatial Understanding
&lt;/figcaption>
&lt;/figure>
&lt;p>It can also help when you can't understand a Japanese menu by identifying and translating the items into Chinese:&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/spatial_understanding_02.jpg" alt="Identify food and translate its name - Spatial Understanding" data-caption="Identify food and translate its name - Spatial Understanding" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Identify food and translate its name - Spatial Understanding
&lt;/figcaption>
&lt;/figure>
&lt;p>Or find stains and teach me how to clean them:&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/spatial_understanding_03.jpg" alt="Detect stains and teach how to clean them - Spatial Understanding" data-caption="Detect stains and teach how to clean them - Spatial Understanding" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Detect stains and teach how to clean them - Spatial Understanding
&lt;/figcaption>
&lt;/figure>
&lt;p>Check out the official demo video: &lt;a href="https://www.youtube.com/watch?v=-XmoDzDMqj4">https://www.youtube.com/watch?v=-XmoDzDMqj4&lt;/a>&lt;/p>
&lt;p>* The official notes mention: &amp;ldquo;Points and 3d bounding boxes are preliminary model capabilities. Use 2D bounding boxes for higher accuracy.&amp;rdquo;&lt;/p>
&lt;br/>
&lt;h4 id="video-analyzer">Video Analyzer&lt;/h4>
&lt;p>Video Analyzer can analyze video scenes, provide summaries, extract text, search for objects, and more.&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/video_analyzer_01.jpg" alt="Summarize a video and add timestamps - Video Analyzer" data-caption="Summarize a video and add timestamps - Video Analyzer" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Summarize a video and add timestamps - Video Analyzer
&lt;/figcaption>
&lt;/figure>
&lt;br/>
&lt;h4 id="map-explorer">Map Explorer&lt;/h4>
&lt;p>Ask Map Explorer questions about countries, landmarks, or geography, and it will pinpoint the answers on Google Maps, making it easy to explore the world ~&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/map_explorer_01.jpg" alt="Find out the tallest mountain in Taiwan - Map Explorer" data-caption="Find out the tallest mountain in Taiwan - Map Explorer" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Find out the tallest mountain in Taiwan - Map Explorer
&lt;/figcaption>
&lt;/figure>
&lt;br/>
&lt;h3 id="other-tests">Other tests&lt;/h3>
&lt;h4 id="recognizing-song-lyrics">Recognizing Song Lyrics&lt;/h4>
&lt;p>Try out the Gemini 2.0 Flash to test its ability to read audio files and recognize song lyrics. Simply provide it with an audio file of a song, and it will identify and organize the lyrics into an LRC format.&lt;/p>
&lt;p>I tested it with Jay Chou's &amp;ldquo;Waiting for You&amp;rdquo; (等你下課) to see how it handles slightly unclear pronunciation. Here's how it performed:&lt;/p>
&lt;p>(Temperature set to 0.5)&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/lyric_gemini_2.0_flash_experimental.jpg" alt="Gemini 2.0 Flash Experimental lyrics recognition results" data-caption="Gemini 2.0 Flash Experimental lyrics recognition results" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Gemini 2.0 Flash Experimental lyrics recognition results
&lt;/figcaption>
&lt;/figure>
&lt;p>Let's also compare with the Gemini 1.5 Pro model:&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/lyric_gemini_1.5_pro.jpg" alt="Gemini 1.5 Pro lyrics recognition results" data-caption="Gemini 1.5 Pro lyrics recognition results" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Gemini 1.5 Pro lyrics recognition results
&lt;/figcaption>
&lt;/figure>
&lt;br/>
&lt;p>When compared to the original lyrics, although both have small recognition errors, it's clear that 2.0 Flash Experimental is better than 1.5 Pro. Additionally, 1.5 Pro mistakenly switched to Simplified Chinese toward the end.&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/lyrics_comparison.jpg" alt="Compare the lyrics recognition results" data-caption="Compare the lyrics recognition results" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Compare the lyrics recognition results
&lt;/figcaption>
&lt;/figure>
&lt;p>Although comparing Flash and Pro models isn't entirely fair, it does show how much Gemini 2.0 has improved.&lt;/p>
&lt;p>However, I don't know why the output lyrics were not complete, stopping about halfway, even though it shouldn't have exceeded the Output length limit of 8192 tokens.&lt;/p>
&lt;br/>
&lt;h4 id="native-tool-use">Native tool use&lt;/h4>
&lt;p>Gemini 2.0 natively supports tools like running code or performing Google searches, allowing for real-time interaction and feedback.&lt;/p>
&lt;br/>
&lt;p>Using Google search can help ensure the accuracy of the answers. It allows the system to gather information from multiple sources and combine them to make the responses more comprehensive.&lt;br />
(The official documentation also mentions that this approach can &amp;ldquo;increase traffic to the source websites&amp;rdquo; 😆)&lt;/p>
&lt;p>For example, if I directly ask it, &amp;ldquo;Which team won the 2024 World Baseball Classic?&amp;rdquo; it might reply that the game hasn't started or that it doesn't know:&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_grounding_unenabled.jpg" alt="Google Search is not enabled" data-caption="Google Search is not enabled" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
Google Search is not enabled
&lt;/figcaption>
&lt;/figure>
&lt;p>However, when the Google Search (Grounding) feature is enabled, it will search the web for information, providing accurate and real-time answers:&lt;/p>
&lt;p>* You can refer to the official documentation for instructions on how to use the program. There's no need to integrate a separate search API: &lt;a href="https://ai.google.dev/gemini-api/docs/grounding" target="_blank" rel="noopener">
Grounding with Google Search
&lt;/a>&lt;/p>
&lt;figure >
&lt;img data-src="https://res.cloudinary.com/jiablog/gemini_2_flash/gemini_grounding_enabled.jpg" alt="After enabling Google Search, it provides the correct answer" data-caption="After enabling Google Search, it provides the correct answer" src="data:image/svg+xml,%0A%3Csvg xmlns='http://www.w3.org/2000/svg' width='900px' height='' viewBox='0 0 24 24'%3E%3Cpath fill='none' d='M0 0h24v24H0V0z'/%3E%3Cpath fill='%23aaa' d='M19 3H5c-1.1 0-2 .9-2 2v14c0 1.1.9 2 2 2h14c1.1 0 2-.9 2-2V5c0-1.1-.9-2-2-2zm-1 16H6c-.55 0-1-.45-1-1V6c0-.55.45-1 1-1h12c.55 0 1 .45 1 1v12c0 .55-.45 1-1 1zm-4.44-6.19l-2.35 3.02-1.56-1.88c-.2-.25-.58-.24-.78.01l-1.74 2.23c-.26.33-.02.81.39.81h8.98c.41 0 .65-.47.4-.8l-2.55-3.39c-.19-.26-.59-.26-.79 0z'/%3E%3C/svg%3E" class="lazyload" style="width:900px;height:;"/>
&lt;figcaption style="text-align: center;">
After enabling Google Search, it provides the correct answer
&lt;/figcaption>
&lt;/figure>
&lt;br/>
&lt;p>The official demonstration shows how to draw charts with code. You can watch this video:&lt;/p>
&lt;iframe width="672" height="378" src="https://www.youtube.com/embed/EVzeutiojWs?si=3hdws4HsnkLkVraO" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;br/>
&lt;br/>
&lt;h4 id="native-image-output">Native image output&lt;/h4>
&lt;p>It also has powerful image generation and editing capabilities. It can precisely modify specific areas of an image without affecting other parts. The commands are more conversational and user-friendly. It can also merge two images or infer possible scenes from existing ones.&lt;/p>
&lt;p>This feature seems impressive and could make AI-powered image editing more practical!&lt;/p>
&lt;p>This part doesn't seem to be available to the public yet, but you can get a glimpse through the official demo videos:&lt;/p>
&lt;iframe width="672" height="378" src="https://www.youtube.com/embed/7RqFLp0TqV0?si=eg8p_AzHp1VSsZLH" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen>&lt;/iframe>
&lt;br/>
&lt;br/>
&lt;h3 id="official-sample-code">Official Sample Code&lt;/h3>
&lt;p>For developers, Google also provides some example code for various features as a starting point for reference:&lt;br />
&lt;a href="https://github.com/google-gemini/cookbook/tree/main/gemini-2">https://github.com/google-gemini/cookbook/tree/main/gemini-2&lt;/a>&lt;/p>
&lt;br/>
&lt;!--adsense-->
&lt;br/>
&lt;h2 id="conclusion">Conclusion&lt;/h2>
&lt;p>Since the model training data is primarily in English, if your results on Gemini aren't satisfactory, it's recommended to use English for your commands.&lt;/p>
&lt;p>All the applications above are based on the Gemini 2.0 Flash model. Once the Gemini 2.0 Pro model is released, the performance will certainly be even better and more powerful~&lt;/p>
&lt;br/>
&lt;p>Below are links to other Gemini 2.0 reference articles. If you&amp;rsquo;re interested, you can click to read further.&lt;/p>
&lt;br/>
&lt;p>If you're interested in Generative AI, make sure to follow the &amp;ldquo;&lt;a href="https://www.facebook.com/jiatool" target="_blank" rel="noopener">
IT Space
&lt;/a>&amp;rdquo; Facebook page to stay updated on the latest posts! 🔔&lt;/p>
&lt;br/>
&lt;br/>
&lt;hr />
&lt;p>References:&lt;br />
&lt;a href="https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/" target="_blank" rel="noopener">
Introduction to Gemini 2.0
&lt;/a>&lt;br />
&lt;a href="https://developers.googleblog.com/en/the-next-chapter-of-the-gemini-era-for-developers/" target="_blank" rel="noopener">
Overview of New Features in Gemini 2.0 Flash
&lt;/a>&lt;br />
&lt;a href="https://www.inside.com.tw/article/37031-google-gemini-2-launch" target="_blank" rel="noopener">
Google 推出新一代 Gemini 2.0！可直接使用搜尋、各種模態無縫融合
&lt;/a>&lt;br />
&lt;a href="https://www.inside.com.tw/article/37032" target="_blank" rel="noopener">
Google：AI 代理時代降臨！一口氣發表自動瀏覽網站、網購、打電動的 AI 助理
&lt;/a>&lt;/p>
&lt;br/>
&lt;blockquote>
&lt;p>Don't be afraid to think different and challenge the status.&lt;/p>
&lt;p align="right">—— Jensen Huang (president and chief executive officer (CEO) of Nvidia)&lt;/p>
&lt;/blockquote></content:encoded><dc:creator>Jia</dc:creator><media:content url="https://blog.jiatool.comimages/cover/gemini_2_flash.jpg" medium="image"><media:title type="html">featured image</media:title></media:content><media:content url="https://blog.jiatool.comimages/posts/gemini_2_flash_meta.jpg" medium="image"><media:title type="html">meta image</media:title></media:content><category>Gemini</category><category>LLM</category><category>AI</category><category>Google</category><category>Share</category></item></channel></rss>