ChatGPT-生成式-AI-实用指南---第2版-早期预览--全-
ChatGPT 生成式 AI 实用指南 - 第2版(早期预览)(全)
原文:Practical Generative AI with ChatGPT - 2nd Edition (Early Access)
译者:飞龙
ChatGPT 生成式 AI 实用指南 - 第2版(早期预览)


-
封面
-
目录
欢迎来到 Packt 早期访问。在本书上市之前,我们为您提供独家预览。撰写一本书可能需要几个月,但我们的作者今天就有尖端信息可以与您分享。早期访问通过提供章节草稿,让您洞察最新进展。目前这些章节的内容可能还些粗糙,但作者将随着时间的推移对其更新。
您可以随时翻读这本书,也可以从头到尾学习;早期访问的设计非常灵活。我们希望您喜欢了解更多关于编写 Packt 书的过程。
-
第 1 章:生成式 AI 简介
-
第 2 章:介绍 OpenAI 和 ChatGPT
-
第 3 章:从提示词设计到提示词工程
-
第 4 章:使用 ChatGPT 提升日常生产力
-
第 5 章:使用 ChatGPT 开发未来
-
第 6 章:使用 ChatGPT 掌握营销
-
第 7 章:使用 ChatGPT 重塑研究
-
第 8 章:使用 ChatGPT 释放视觉创造力
-
第 9 章:探索 GPTs
-
第 10 章:利用 OpenAI 模型及其 API 实现企业级应用
-
第 11 章:趋势用案例
-
第 12 章:结语与最后思考
1 生成式 AI 简介
在 Discord 上加入我们的书籍社区

你好!欢迎来到ChatGPT 和 OpenAI 终极指南!在本书书中,我们将探索迷人的生成式 人工智能(GAI)世界及其突破性的应用。生成式 AI 改变了我们与机器交互的方式,使计算机能够在没有显式人类指令的情况下创建、预测和学习。自 2023 年 11 月 ChatGPT 发布以来,我们见证了自然语言处理、图像和视频合成以及许多其他领域前所未有的进步。无论您是好奇的初学者还是经验丰富的从业者,本指南都将为您提供知识和技能,帮助您在激动生成式 AI 领域中穿行。因此,让我们深入研究,从我们所处环境的一些定义开始。
本章提供了生成式 AI 领域的概述,该领域由使用机器学习(ML)算法创建新的独特数据或内容组成。
它重点关注生成式 AI 在各个领域的应用,如图像合成、文本生成和音乐创作,强调了生成式 AI 变革各行业的潜力。这一关于生成式 AI 的介绍将为该技术的所在位置提供背景,以及将其置于广阔的 AI、ML 和深度学习(DL)世界中的知识。然后,我们将通过具体的示例和近期进展,深入探讨生成式 AI 的主要应用领域,让您可以熟悉它对企业和整个社会可能产生的影响。
此外,了解通往当前生成式 AI 领先水平的研究历程,将让您能够更好地理解近期进展和顶尖模型的基础。
对于所有这些内容,我们将涵盖以下主题:
-
理解生成式 AI
-
探索生成式 AI 的领域
-
生成式 AI 研究的历史和现状
在本章结束时,您将熟悉生成式 AI 这一激动的世界、它的应用、背后的研究历史以及现状,这些发展可能已经且目前正在对企业产生颠覆性的影响。
介绍生成式 AI
AI 近年来一直取得了显著进展,其中增长相当大的领域之一就是生成式 AI。生成式 AI 是 AI 和 DL 的子领域,侧重于使用 ML 技术在现有数据上训练的算法和模型来生成新内容,如图像、文本、音乐和视频。
为了更好地理解 AI、ML、DL 和生成式 AI 之间的关系,请将 AI 视为基础,而 ML、DL 和生成式 AI 代表了日益专业和专注的研究与应用领域:
-
AI 代表创建可以执行任务的系统的广泛领域,展示出人类的智能和能力,并能够与生态系统交互。
-
ML 是一个分支,专注于创建算法和模型,使这些系统能够随着时间和训练而进行自我学习和改进。ML 模型从现有数据中学习,并随着增长自动更新其参数。
-
DL 是 ML 的一个子分支,因为它包含了深度 ML 模型。这些深度模型被称为
神经网络,特别适用于计算机视觉或自然语言处理(NLP)等领域。当我们谈论 ML 和 DL 模型时,通常指的是判别式模型,其目标是在数据之上进行预测或推断模式。 -
最后,我们得到了生成式 AI,它是 DL 的进一步子分支,它不使用深度人工神经网络对现有数据进行聚类、分类或预测:它使用那些强大的人工神经网络模型来生成全新的内容,从图像到自然语言,从音乐到视频。
接下来的图表显示了这些研究领域之间的相互关系:

图 1.1 – AI、ML、DL 和生成式 AI 的关系
生成式 AI 模型可以在海量数据上进行训练,然后它们可以根据数据中的模式生成新的示例。这种生成过程与判别式模型不同,后者的目标是预测给定示例的类别或标签。
具有生成式 AI 特性的神经网络类型被称为大基础模型(LFMs),对于 ChatGPT 这样的语言模型,我们称之为大语言模型(LLMs)。
大语言模型是一种以 Transformer 架构为特征的人工神经网络。它们的特点是具有海量的参数(以万亿计),并已在数十亿词汇上进行过训练。给定训练集,LLMs 能够理解并生成自然语言。
尽管文本理解和生成可能是生成式 AI 最突出的特征之一,但该领域涵盖了许多领域。
生成式 AI 的领域
近年来,生成式 AI 取得了显著进展,并将其应用扩展到广泛的领域,如艺术、音乐、时尚、建筑等。在某些领域中,它确实改变了我们创造、设计和理解周围世界的方式。在其他领域,它正在改进并使现有的流程和操作更加高效。
生成式 AI 被应用于许多领域这一事实意味着它的模型可以处理不同类型的数据,从自然语言到音频或图像。让我们了解生成式 AI 如何处理不同类型的数据和领域。
文本生成
生成式 AI 的最大应用之一(也是我们在书中涵盖最多的部分)是它在自然语言中产生新内容的能力。事实上,生成式 AI 算法可以用于生成新的文本,如文章、诗歌和产品描述。
例如,由 OpenAI 开发的 GPT-4o 等语言模型可以在海量文本数据上进行训练,然后用于生成不同语言(包括输入和输出)的新的、连贯且符合语法的文本,并从文本中提取相关特征,如关键词、主题或全文摘要。
以下是一个使用 GPT-3 的示例:

图 1.2 – ChatGPT 响应用户提示的示例,并添加了引用
接下来,我们将转向图像生成。
图像生成
生成式 AI 在图像合成领域最早且最为人知的例子之一,是 I. Goodfellow 等人在 2014 年的论文《生成对抗网络》(Generative Adversarial Networks)中引入的生成对抗网络(GAN)架构。GAN 的目的是生成与真实图像无区分的真实图像。这种能力具有许多有趣的业务应用,例如为训练计算机视觉模型生成合成数据集、生成真实的商品图像,以及为虚拟现实和增强现实应用生成真实的图像。
以下是一些并不存在的人脸示例,因为它们完全是由 AI 生成的:

图 1.3 – 在 https://this-person-does-not-exist.com/en 由 GAN StyleGAN2 生成的虚构脸庞
然后,2021 年,OpenAI 在该领域引入了新的生成式 AI 模型 DALL-E。与 GANs 不同,DALL-E 模型设计用于根据自然语言描述生成图像(GANs 以随机噪声向量输入),并可以生成广泛的图像,这些图像可能看起来并不真实,但仍描绘所需的概念。
DALL-E 在广告、产品设计和时尚等创意产业中具有巨大的潜力,可以创建独特且创意的图像。
自首次发布以来(2024 年 12 月),DALL-E 有了巨大的改进,如下例所示。让我们看看下文由 DALL-E 在其诞生之初创作的一件艺术作品:

图 1.4 – 以自然语言提示作为输入 DALL-E 生成的图像
现在让我们看看 DALL-E3(本书编写时的模型最新版本)可以产生什么(在这里我使用了由 DALL-E3 提供的 Microsoft Image Creator。你可以在 copilot.microsoft.com/images/create 尝试它):

图 1.5 – 以自然语言提示作为输入 DALL-E3 生成的图像
看到这个模型在不到 18 个月内的改进水平令人印象深刻。有趣的是,我们仅仅接触到了过去几个月生成式 AI 领域发生的巨大改进的冰角。
音乐生成
音乐生成的生成式 AI 早期方法可以追溯到 50 年代,源于算法作曲领域的研究,这是一种使用算法生成音乐作品的技术。事实上,1957 年,Lejaren Hiller 和 Leonard Isaacson 创建了《伊利亚克弦四奏组曲》(Illiac Suite for String Quartet)(www.youtube.com/watch?v=n0njBFLQSk8,这是第一首完全由 AI 作曲的音乐。以来,音乐生成式 AI 一直是几十年来持续的研究课题。在近年的发展中,新的架构和框架已经在公众中普及,例如 Google 在 2016 年引入的 WaveNet 架构,它能够生成高质量的音频样本;或者同样由 Google 开发的 Magenta 项目,它使用循环神经网络(RNNs)和其他机器学习技术来生成音乐和其他艺术形式。然后,在 2020 年,OpenAI 也宣布了 Jukebox,这是一种生成音乐的神经网络,可以在音乐和人声风格、流派、参考艺术家等方面自定义输出。
这些和其他框架构成了许多用于音乐生成的 AI 作曲助手的基础。例如是由 Sony CSL Research 开发的 Flow Machines。这个生成式 AI 系统在大型音乐作品数据库上进行训练,以创建各种风格的新音乐。它被法国作曲家 Benoît Carré 用于创作一张名为 Hello World 的专辑(www.helloworldalbum.net/),该专辑包含了与多位人类音乐家的合作。
在这里,你可以看到完全由 Music Transformer 生成的曲目示例,它是 Magenta 项目中的模型之一:

图 1.6 – Music Transformer 允许用户收听 AI 生成的音乐表演
音乐领域中生成式 AI 的另一个惊人应用是语音合成。确实可以找到许多 AI 工具,可以根据文本输入以知名歌手的声音创建音频。
例如,如果你一直想知道你的歌曲由凯尼·韦斯特演唱会什么样,现在可以使用 FakeYou.com([fakeyou.com/]( https://fakeyou.com/ ))或 UberDuck.ai(https://uberduck.ai/
图 1.7 – 使用 UberDuck.ai 进行语音合成
我必须说,结果非常令人印象深刻。如果你想玩玩,还可以尝试你所有喜爱的卡通声音,例如维尼的浦……让我们更进一步。如果我们能从零开始生成一首歌,只需让生成式 AI 这样做会怎样?嗯,现在我们可以无缝地做到,而无需任何音乐知识。在音乐市场兴起的生成式 AI 产品中,一个很棒例子是 Suno,它的使命是“……我们正在建立一个任何人都能创作绝佳音乐的未来。无论你是浴室歌手还是榜单艺术家,我们都打破了你与梦想歌曲之间的障碍。不需要乐器,只需要想象力。”(来源:suno.com/about)。

你能相信这成了我 2024 年夏季的热单吗?如果你也想创作自己的夏季热门单,可以在 suno.com/create 免费尝试。
视频生成
视频生成的生成式 AI 在发展时间线上与图像生成相似。事实上,视频生成领域的一个关键进展之一是生成对抗网络的发展的发展。得益于它们在生成真实图像方面的准确性,研究人员已开始将这些技术应用于视频生成。基于 GAN 的视频生成的著名例子之一是 DeepMind 的 Motion to Video,它从单张图像生成高质量视频。另一个很棒的例子是 NVIDIA 的视频转视频合成(Vid2Vid)深度学习框架,它使用 GANs 从输入视频合成高质量视频。
Vid2Vid 系统可以生成时间一致的视频,这意味着它们在时间推移过程中保持平滑且真实的运动。该技术可以执行各种视频合成任务,例如:
-
将视频转换为另一个领域(例如将白天视频转换为夜晚视频,或将草图转换为真实的图像)
-
修改现有视频(例如,更改视频中对象的风格或外观)
-
从静态图像创建新视频(例如,使一系列静态图像动画化
2022 年 9 月,Meta 的研究人员宣布了 Make-A-Video (makeavideo.studio/) 的开放可用,这是一种新的 AI 系统,允许用户将自然语言提示转换为视频片段。在这些技术背后,你可以识别出我们目前为止在其他领域提到的许多模型:用于提示的语言理解、结合图像生成实现的图像和运动生成,以及由 AI 作曲家制作背景音乐。
现在,我们上面提到的所有内容与最新的文本视频模型相比都相形见猔。例如,OpenAI 在 2024 年 2 月宣布了一种名为 SORA 的新型文本视频模型,并发布了一些出乎意料的早期实验:


图 1:根据自然语言提示由 SORA 生成的视频。来源:https://openai.com/index/sora/
我非常建议你访问 SORA 网页查看它创建的惊人视频。在编写此时,SORA 尚未公开可用,因为它正在接受 OpenAI 红队运行的多项测试。
总的来说,生成式 AI 多年来影响了许多领域,某些 AI 工具已经能够持续支持艺术家、机构和普通用户。未来看起来非常有前景;然而,在跳转到目前市场上最顶尖的模型之前,我们首先需要深入了解生成式 AI 的根源、其研究历史以及最终导致当前 OpenAI 模型的最新进展。
研究的历史与现状
在之前的章节中,我们概述了生成式 AI 领域最新、最前沿的技术,这些都是在几年内开发的。然而,该领域的研究可以追溯到几十年前。
我们可以将生成式 AI 领域研究的起点标记在 20 世纪 600 年代,当时 Joseph Weizenbaum 开发了聊天机器人 ELIZA,这是 NLP(自然语言处理)系统的早期示例之一。它是一个基于规则的简单交互系统,旨在根据文本输入的响应来娱乐用户,为 NLP 和生成式 AI 的进一步发展铺平了道路。然而,我们知道现代生成式 AI 是深度学习(DL)的一个子领域,尽管第一个人工神经网络(ANNs)早在 20 世纪 40 年代就被引入,但研究人员面临着许多挑战,包括计算能力有限以及对大脑生物基础缺乏了解。因此,ANNs 直到 20 世纪 80 年代才获得广泛关注,当时除了新硬件和神经科学的发展外,backpropagation(反向传播)算法的出现也促进了 ANNs 的训练阶段。事实上,在反向传播出现之前,训练神经网络是非常困难的,因为无法高效地计算误差相对于每个神经元的参数或权重的梯度,而 backpropagation 使得自动化训练过程成为可能,并实现了 ANNs 的应用。
随后,到 2000 年代和 2010 年代,计算能力的提升加上海量的可用训练数据使得深度学习更具实用性并对公众开放,随之推动了研究的增长。
2013 年,Kingma 和 Welling 在他们的论文 Auto-Encoding Variational Bayes 中引入了一种新的模型架构,被称为变分自动编码器(VAEs)。VAEs 是基于变分推理概念的生成模型。它们提供了一种通过数据的紧凑表示进行学习的方法,将数据编码到低维空间(称为潜空间,通过 encoder 编码器组件),然后将其解码回原始数据空间(通过 decoder 解码器组件)。
VAEs 的关键创新在于引入了潜空间的概率解释。encoder 不再学习输入到潜空间的确定性映射,而是将输入映射到潜空间上的概率分布。这使得 VAEs 可以通过从潜空间中采样并将样本解码到输入空间来生成新样本。
例如,假设我们想训练一个 VAE,它可以创建看起来像真实的猫和狗的新图片。
为了做到这一点,VAE 首先接收一张猫或狗的照片,将其压缩为潜空间中一组较小的数字,这些数字代表了图片最重要的特征。这些数字被称为潜变量。
然后,VAE 获取这些潜变量并使用它们创建一张看起来像真实的猫或狗的照片。这张新图片可能与原始图片有些不同之处,但它看起来应该属于同一组图片。
VAE 随着时间的推移,通过将其生成的图片与真实图片进行比较并调整其潜变量,使生成的图片看起来更真实,从而进步。
VAEs 为生成式 AI 领域的快速发展铺平了道路。事实上,仅 1 年后,Ian Goodfellow 就引入了 GANs(生成对抗网络)。与主要元素为 encoder 和 decoder 的 VAEs 架构不同,GANs 由两个组成——generator(生成器)和 discriminator(判别器)——它们在了一个零和博弈中相互对抗。
generator 创建伪数据(对于图像而言,它创建一张新图像),目的是看起来像真实数据(例如,猫的照片)。discriminator 接收真实数据和伪数据,并试图区分它们——在我们艺术伪造者的例子中它就是 critic(评论家)。
在训练期间,generator 尝试创建能够欺骗 discriminator 认为其为真实的数据,而 discriminator 则试图变得更好地区分真伪数据。这两个部分在一个被称为对抗训练(adversarial training)的过程中被共同训练。
随着时间的推移,generator越来越擅长创建看起来像真实数据的伪数据,而 discriminator 越来越擅长区分真伪数据。最终,generator 创建伪数据的能力变得非常强大,以至于即使是 discriminator 也无法区分真伪数据。
这是一个完全由 GAN 生成的人脸示例:

图 1.8 – GAN 生成的写实脸(取自 Progressive Growing of GANs for Improved Quality, Stability, and Variation, 2017: https://arxiv.org/pdf/1710.10196.pdf )
两种模型——VAEs 和 GANs 的目的是生成与原始样本无法区分的新数据,而且自诞生以来,它们的架构不断改进,并伴随着新模型如 Van den Oord 团队提出的 PixelCNNs 以及 Google DeepMind 开发的 WaveNet 的发展,推动了音频和语音生成的进步。
另一个伟大的里程碑是 2017 年,当时 Google 研究人员在论文中引入了一种名为 Transformer 的新架构——Attention Is All You Need 是 Google 研究人员在论文中引入的。它在语言生成领域具有革命性,因为它在允许并行处理的同时保留了语言上下文的记忆,性能优于此前基于 RNN 或 长短期记忆(LSTM)框架的语言模型尝试。
Transformers 确实是名为 Bidirectional Encoder Representations from Transformers (BERT) 的大规模语言模型的基础,该模型由 Google 于 2018 年引入,并迅速成为 NLP 实验的基准。
Transformers 也是 OpenAI 引入的所有 Generative Pre-Trained (GPT) 模型的基础,包括 ChatGPT 背后的模型 GPT-3,以及其他大型基础模型(Large Foundation Models)。
尽管在这些几年取得了大量的研究和成就,但直到 2022 年下半年,公众的注意力才开始转向生成式 AI 领域。
并非巧合,2022 年被称为“生成式 AI 之年”。这一年,强大的 AI 模型和工具在普通民众中变得普及:基于扩散的图像服务(MidJourney、DALL-E 2 和 Stable Diffusion)、OpenAI 的 ChatGPT、文本生成视频(Make-a-Video 和 Imagen Video)以及文本生成 3D 工具(DreamFusion、Magic3D 和 Get3D)都向个人用户开放,有时甚至是免费的。
这产生了颠覆性的影响,原因有二:
-
一旦生成式 AI 模型在公众中普及,每个个人用户或组织都有可能尝试并欣赏其潜力,即使他们不是数据科学家或 ML 工程师。
-
这些新模型的输出及其嵌入的创造力在客观上是令人惊叹的,而且往往令人担忧。一种针对个人和政府的、适应变革的紧迫呼声随之而起。
因此,在不久的将来,我们可能会见证个人使用和企业级项目采用 AI 系统的激增。
18 个月后:主要趋势与创新
从 2022 年 11 月到今天,我们见证了生成式 AI(GenAI)领域海量的创新。这些创新许多与向公众开发并发布的全新模型有关,例如 OpenAI 的 GPT-4o 和 DALL-E3,以及 Google 的 Gemini、Meta 的 Llama 3、微软的 Phi3 等。
然而,最引注目的成就可能在于我们与这些模型交互以及围绕它们构建应用程序的方式。在本节中,我们将探索标记了生成式 AI 应用最流行的参考架构的三大主要进展。
检索增强生成 (Retrieval Augmented Generation)
ChatGPT 以及广义上的 LLMs 的早期局限性之一是知识库截止。事实上,LLMs 的知识局限于它们训练过的训练集(由互联网上的公共信息组成),只要这些信息不是详尽的,它们就是过时的。此外,它们还缺少可能与我们或我们的组织相关的私有知识库。例如,如果你问 ChatGPT “我们公司员工健康保险的政策是什么?”,模型将无法回答,因为它无法访问这些信息。
为了绕过这一局限,设计了一种新框架,允许 LLMs 导航我们提供的自定义文档。这个框架被称为检索增强生成(RAG)。
RAG 背后的思想是将 LLM 与我们想要导航的知识库解耦。为了实现这一点,外部知识库需要通过一个名为 embedding(嵌入)的过程转换为数值向量,并存储在专门的 Vector Database(向量数据库)中。
嵌入是在低维空间(如向量)中表示高维非数值数据(如单词或句子)的方法。文本嵌入可以捕捉文本的语义和语法特征,例如含义、上下文和相似性。
每个嵌入都是一个浮点数向量,使得向量空间中两个嵌入之间的距离与原始格式中两个输入之间的语义相似性相关。
例如,如果两个概念相似,那么它们的向量表示也相似的。

RAG 由三个阶段组成:
- 检索 (Retrieval):给定用户的查询及其对应的向量,检索出最相似的文档片段(那些对应的向量接近用户查询向量的片段),并将其作为
LLM的基础上下文。

- 增强 (Augmentation):通过额外的指令、规则、安全护栏以及类似于提示词工程(prompt engineering)技术的典型实践来丰富检索到的上下文(我们将在第 3 章中介绍提示词工程主题)。

- 生成 (Generation):基于增强的上下文,
LLM生成对用户查询的响应。

整个流程如下:

RAG 结合了生成式模型和信息检索系统的优势,以增强生成内容的质量和相关性。传统的生成式模型仅依赖其训练数据来产生响应,这有时会导致过时或无关的信息。RAG 通过在生成过程中整合外部知识库解决了这一局限性。
例如,OpenAI 的 ChatGPT 在通过 RAG 增强后,可以从可靠源获取当前数据,以回答关于近期事件的查询,这本属于其训练截止的范围。这种检索与生成的协同扩大了生成式 AI 的适用性,使其在动态、信息丰富的环境中更加稳健和实用。
多模态性 (Multimodality)
在本章的第一段,我们看到了生成式 AI 的各个领域,从文本到图像,从视频到音乐。通常,大型基础模型倾向于特定领域,正如我们在语言理解和生成领域的大型语言模型(LLMs),或图像生成领域的 DALL-E3 中看到的那样。
然而,生成式 AI 的最新进展使得大型多模态模型(LMMs)成为,它们可以处理和生成不同类型的数据,如文本、图像、音频和视频。
LMMs 与“标准”大型语言模型(LLMs)共享大型基础模型典型的泛化和适应能力。然而,LMMs 能够处理异构数据,其理念是镜像人类与周围生态系统交互的方式——即通过我们所有的感官。
多模态模型的一个很好的例子是 OpenAI 的 GPT-4o,它能够通过文本、图像和音频与用户交互。让我们来看例子:

如你所示,模型能够分析图像并对其进行推理。现在让我们让模型生成一张插图:

关于 LMMs 最有趣的事实是,它们保留了推理能力,使得它们适用于异构数据语境下的复杂推理。让我们考虑最后这个例子(仅显示响应的前几行):

正如你所想象的,这为各个行业开启了应用前景,我们将在后续章节中看到一些示例。
AI 代理 (AI Agents)
在之前的节中,我们揭示了 LFMs 在共鸣和生成内容方面表现出色。然而,它们缺乏一种能力,即超越单个用户采取行动并与周围生态系统交互的能力。例如,如果我们希望 LLM 不仅能生成一条出色的 LinkedIn 帖子,还能将其发布在我们的页面上,呢?
为了克服这一局限性,AI 代理(AI agents)成为了关键角色。但它们究竟是什么呢?代理可以被看作是由大语言模型(LLM)驱动的 AI 系统,根据用户的查询,它们能够在我们允许范围内与周围的生态系统进行交互。生态系统的边界由我们为代理提供的工具(或插件)决定(在我们之前的示例中,我们可能会为代理提供一个 LinkedIn 插件,以便它能够发布生成的内容)。
代理体由以下要素组成:
-
一个作为 AI 系统推理引擎的 LLM(大语言模型)。
-
一条系统消息(system message),用于指示代理以特定的方式行为和思考。例如,你可以将一个代理体设计为学生的助教,并带有如下系统消息:“你是一个助教。给定学生的查询,永远提供最终答案,而是提供一些提示来帮助他们获得答案”。
-
一组代理可以利用其与周围生态系统进行交互的工具。
AI 代理体是“LLM 作为应用程序推理引擎”这一含义的完美体现。事实上,代理体的优之处在于,它们可以选择最合适的工具来完成用户的请求。例如,假设我们有一个用于生成 LinkedIn 内容的 AI 代理,并为其提供了两个工具:一个 LinkedIn 插件和一个网络搜索插件(每个插件都对其功能有适当的描述)。然后,让我们探索代理面对两个不同问题时的行为:
-
生成一个关于小狗在山间散步的故事🡪代理将在不涉及任何插件的情况下生成故事。
-
生成一个关于米兰当前天气的故事🡪代理将调用网络搜索插件获取米兰当前的天气。
-
生成一个关于米兰当前天气的 LinkedIn 帖子并发布在我的个人页上:代理将调用网络搜索插件获取米兰当前的天气,并调用 LinkedIn 插件将其发布到我的个人页。
指令和插件集的组合使得 AI 代理体极其多功能,你可以创建高度特化的实体来应对特定场景。
而且不仅于此。
既然你可以创建一个相互对话和协作的代理体团队,为什么要只用一个代理呢?想象一下多个代理,每个都有特定的专业知识和目标,它们之间相互通信和交互以完成任务。这就是多代理应用(multi-agent applications)的模样,在过去的几个月里,这种模式开始表现出涌现行为。
让我们考虑以下示例。我们想生成一个关于气候变化的电梯演讲(elevator pitch)。我们需要更新信息来完成(最新趋势和研究、未来展望等),以及基于学术论文的坚实研究。此外,我们需要保持简洁、敏锐且高效,在非常短的演讲中交付所有关键信息。
现在,我们可以向单个代理提出所有要求,为它提供所有需要的工具和长长的指令来完成任务。然而,事实证明,当我们给 LLM 安排“太多的要做的事情”时,它们的表现往往会变差。相反,让我们使用多代理方法,创建一个由以下 AI 专业人士组成的团队:
-
一个市场分析师,可以在网络上搜索有关气候变化的最新新闻;这将是一个拥有网络搜索插件和特定搜索指令的代理;
-
一个专家研究员,可以轻松查阅关于气候变化的学术论文;这将是一个拥有
Arxiv插件和如何检索相关信息特定指令的代理; -
一个演讲专家,可以轻松地将所有信息整合进一个电梯演讲中;这将是一个拥有如何进行完美演讲的恰当指令的代理;
-
一个批评者,将审查演讲并在需要时向演讲专家提出修改建议;这将是一个拥有如何进行完美演讲的恰当指令的代理;
因此,当用户提问“生成一个关于当前气候变化问题的电梯演讲”时,所有代理都可以开始开展该项目。
有很多框架可以帮助开发者构建多代理应用(包括 AutoGen、LangGraph、CrewAI),特别是涉及到我们希望代理遵循的“流程”时。例如,我们可能希望强制特定的迭代次数;或者所有代理至少被调用一次;或者甚至让我们作为用户参与每次迭代,提供进一步的反馈并将其整合到下次迭代中。
在编写此时,多代理框架显示出广阔的进展,但它仍然处于实验阶段,可能远未达到企业级水平。尽管如此,它是 LLM 后方卓越推理能力的一种瞥,展示了它们如何开启解决问题和创造的新方法。
小语言模型
大语言模型名如其义非常大。这意味着包含 LLM 的人工神经网络架构由海量参数组成,数量级达数万亿。理想情况下,参数数量越高,LLM 的推理能力越强。然而,高参数量也意味着高昂的训练和托管成本,因为需要强大的 AI 基础设施。此外,这些模型消耗的能源引发了关于 LLM 训练对环境的影响及其长期可持续性的严肃问题。为了让你有一个直观的概念,以下是与不同规模(以参数数量衡量)的开源 LLM 相关的一些数字:

图 2:在同一数据中心训练不同模型所需的碳足迹和算力。来源:https://arxiv.org/pdf/2302.13971
幸运的是,在过去的几个月里,我们见证了对小模型研究兴趣和投入的增加,目标是在保持良好推理能力的同时减少参数数量。这种理念是,小模型即使“智能程度较低”,但如果专注于特定的活动,仍然可以满足需求。
这些较小的模型被称为小语言模型(SMLs),除了对基础设施的要求更轻、更低之外,它们还表现出了惊人的性能。
例如,如果我们考虑 Phi-3 的较小版本(微软开发的流行的 SLM 家族的一部分),它仅有 70 亿参数,我们可以看到它在所有最流行的 LLM 基准测试中都胜了 GPT-3.5-turbo:

图 3:Phi-3 与其他 SLM 和 LLM 的能力对比表。来源:https://azure.microsoft.com/en-us/blog/introducing-phi-3-redefining-whats-possible-with-slms/
现在,我们可能会认为 GPT-3.5-turbo 已经过时,但我们必须记住,凭借 1750 亿参数在一年前还是市场上最强大的模型,看到一个 70 亿的模型能够产生更好的结果是非常令人惊讶的。
SLM 领域绝对值得关注的研究方向,特别是当我可能希望在本地部署模型甚至通过微调对其进行定制的场景时(我们将在下一章介绍微调)。
在本章中,我们探索了生成式 AI 的精彩世界及其各个应用领域,包括图像生成、文本生成、音乐生成和视频生成。我们学习了由 OpenAI 训练的 ChatGPT 和 DALL-E 等生成式 AI 模型如何利用 DL 技术学习大型数据集中的模式,并生成既新颖又连贯的新内容。我们还讨论了生成式 AI 的历史、起源以及目前的研究现状。
本章的目标是为生成式 AI 的基础提供坚实的基础,并启发你进一步探索这一迷人的领域。
在下一章中,我们将关注目前市场上最有前景的技术之一 ChatGPT:我们将介绍其背后的研究以及 OpenAI 的开发过程、模型架构,以及截至目前它可以解决的主要用例。
参考文献
2 OpenAI 与 ChatGPT:超越市场炒作
在 Discord 上加入我们的书籍社区

本章介绍了 OpenAI 及其最著名的开发成果——ChatGPT,重点介绍了它的历史、技术和能力。
总目标是深入了解 ChatGPT 如何在各种行业和应用中用于改进通信和自动化流程,最后这些应用将如何影响技术世界及更广泛的领域。
我们将通过以下主题涵盖所有这些内容:
-
什么是 OpenAI?
-
OpenAI 模型家族概览
-
通往 ChatGPT 之路:其背后的模型数学
-
ChatGPT:领先水平
技术要求
为了能够测试本章中的示例,你需要以下内容:
-
一个 OpenAI 账户以访问
Playground和Models API(openai.com/api/login20) -
你喜欢的 IDE 环境,例如
Jupyter或Visual Studio -
安装了
Python3.7.1+ (www.python.org/downloads) -
安装了
pip(pip.pypa.io/en/stable/installation/) -
OpenAI Python 库 (
pypi.org/project/openai/)
什么是 OpenAI?
OpenAI 是由 Elon Musk、Sam Altman、Greg Brockman、Ilya Sutskever、Wojciech Zaremba 和 John Schulman 于 2015 年创立的研究机构。正如 OpenAI 网页上所述,其使命是“确保通用人工智能 (AGI) 造福全人类”。由于其通用性,AGI 旨在具有学习和执行广泛任务的能力,而无需针对特定任务进行编程。
自 2015 年以来,OpenAI 将研究重点放在深度强化学习 (DRL),这是机器学习 (ML) 的子集,结合了强化学习 (RL) 和深度神经网络。该领域的第一个贡献可以追溯到 2016 年,当时公司发布了 OpenAI Gym,这是一个供研究人员开发和测试 RL 算法的工具包。

图 2.1 – Gym 文档首页 (https://www.gymlibrary.dev/)
OpenAI 在该领域持续进行研究和贡献,但它最著名的成就与生成式模型——生成式预训练 Transformer (GPT) 相关。
在本章中,我们将探索第一种方法,并在第 11 章中介绍如何整合 OpenAI 模型的 API。
在进入 Playground 之前,让我们先概览一下 OpenAI 的模型家族。
OpenAI 模型家族
在过去的几年里,OpenAI 在模型开发领域取得了巨大的进展,以惊人的速度发布了新版本的模型。在本节中,我们将看到按领域划分的主要模型。
-
语言模型 🡪 OpenAI 的 GPTs 是先进的语言模型,旨在根据给定的提示词(prompts)生成类人类文本。它们多才且,可以用于各种自然语言处理任务,如文本补全、翻译、摘要和编码。得益于其旗舰模型系列 GPTs,OpenAI 在这一领域展现了卓越的性能。自 2022 年 11 月 ChatGPT 发布以来,OpenAI 发布了:
-
GPT-3.5-turbo,第一版 ChatGPT 背后的模型 -
GPT-4(第一个多模态模型)和GPT-4-turbo(针对聊天和助手进行了优化) -
GPT-4o,最先进的多模态模型 -
GPT-4o mini,GPT-4o的较小版本,速度更快且更便宜。
-
除了这些最新模型外,用户还可以访问一组所谓的“GPT 基础模型”,如 Babbage-002 或 davinci-002(GPT-2),它们代表了旧版的补全 API(pletions APIs)(我们将在下一节介绍补全 API)。
-
图像模型 🡪 OpenAI 的图像模型(如
DALL-E)旨在根据描述生成并处理图像。DALL-E模型可以创建高度详细且富有想象力的视觉效果,使用户创作独特的艺术品、设计概念等。例如,DALL-E 3(该模型的最新版本)能够根据复杂的提示词创建复杂且富有创意的图像,推向 AI 在视觉艺术领域的边界。这些模型在创意产业、数字营销以及任何受益于高质量、定制化视觉效果的领域都特别有用。 -
文本转语音和语音转文本 🡪 在第 1 章中,我们已经提到了 OpenAI 开发的语音转文本(STT)模型
Whisper,它可以处理和分析音频输入。此外,OpenAI 还开发了文本转语音(TTS)模型,提供既清晰又具表现力的高质量输出,使其适用于从客服机器人到教育工具的广泛应用。OpenAI 开发的文本转语音(TTS)模型将文字文本转换为口语语言,产生真实且自然的音频。这些模型对于创建语音助手、提高视障用户的无障碍性以及生成自动公告至关重要。 -
文本转视频 🡪 随着
SORA的发布,OpenAI 展示了其尖端的文本转视频模型,模型能够根据文本描述生成现实且富有想象力的视频场景。通过借鉴DALL-E的技术并结合 Transformer,SORA可以创建长达一分钟的高保真视频。它在保持 3D 一致性、物体持久性和模拟视频内交互方面表现出色。尽管在准确建模复杂物理和确保时间连贯性方面仍面临挑战,但SORA在创意产业具有巨大潜力,为视频制作和故事讲述提供了新的可能性。 -
嵌入(Embeddings) 🡪 OpenAI 的嵌入模型将文本转换为称为向量(或嵌入)的数值表示,这些向量捕捉了语义含义并投影到多维向量空间中。
该空间中不同实例之间的数学距离代表了它们在含义上的相似性。例如,想象单词 queen(皇后)、woman(女性)、king(国王)和 man(男性)。理想情况下,在我们在这个向量空间中,如果表示是正确的,我们希望实现以下效果:

图 2.10 – 单词向量方程示例
这意味着 woman 和 man 之间的距离应该等于 Queen 和 King 之间的距离。
嵌入在智能搜索场景中非常有用。事实上,通过获取用户输入和用户想要搜索的文档的嵌入可以计算输入与文档之间的距离度量(即 cosine similarity 余弦相似度)。通过这种方式,我们可以检索在数学距离上更接近用户输入的文档。
OpenAI 的嵌入模型 text-embedding-ada-002 和 text-embedding-3-large 通过创建文本的稠密向量表示,在文本相似性、文本搜索和代码搜索等任务中提供了顶尖性能。这些嵌入允许对大型文本数据集进行高效有效的比较,提高了搜索的准确性和相关性。
- 审核(Moderation) 🡪 OpenAI 的审核模型旨在检测并过滤文本中不适当、有害或不安全的内容。这些模型通过识别潜在的冒犯或有害语言,对于维护安全尊重的在线环境至关重要。最新的审核模型
text-moderation-007在确保跨平台内容合规安全方面稳健且有效。它帮助开发者和公司执行社区准则并防止有害内容的传播,从而促进更安全的数字交互。
总体来说,OpenAI 近年来覆盖范围越来越广,用顶尖模型解决了生成式 AI 的各个领域。
在 Playground 中尝试 OpenAI 模型
要访问你的 OpenAI Playground,你需要创建一个 OpenAI 账户并访问 platform.openai.com/playground。落地页的样式如下:

图 2.4 – https://platform.openai.com/playground 的 OpenAI Playground
从图 2.4 所示,Playground 提供了一个用户可以在其中与模型交互的界面,你可以在界面右侧选择模型。
在深入研究 Playground 的主要部分之前,让我们先定义一下章中会看到的术语:
-
Token(标记):Token 可以被视为 API 用于处理输入提示词的单词碎片或片段。与完整的单词不同,Token 可能包含尾随空格甚至部分子词。为了更好地理解 Token 的长度概念,有一些通用指南需要记住。例如,英语中的一个 Token 大约相当于四个字符,或四分之三个单词。
-
提示词(Prompt):在自然语言处理(NLP)和机器学习(ML)语境下,提示词是指输入 AI 语言模型以生成响应或输出的文本。提示词可以是问题、陈述或句子,用于为语言模型提供上下文和方向。
-
上下文(Context):在 GPT 领域,上下文是指用户提示词之前的单词和句子。语言模型利用这些上下文根据训练数据中发现的模式和关系生成最可能的下一个单词或短语。
-
模型置信度(Model confidence):模型置信度是指 AI 模型对特定预测或输出分配的确定程度或概率。在 NLP 语境下,模型置信度通常用于表示 AI 模型对其针对输入提示词生成的响应的正确性或相关性有多信心。
函数 (Functions)
函数是对第 1 章中引入的工具或插件概念的另一种定义。通过函数,我们为模型提供了一种额外的技能,模型可以调用该技能来完成用户的任务。函数始终包含一个自然语言编写的描述,以便模型知道何时调用它。
前面的定义对于理解如何使用 Azure OpenAI 模型系列以及如何配置它们的参数至关重要。
在 Playground(游场)中,有四个部分可以与模型交互:
聊天 (Chat)
在这里你可以测试今天可用的所有聊天模型,包括仅文本的模型(如 GPT-3.5)或多模态模型(如 GPT-4o)。你可以提供系统消息(system message)——即提供给模型的指令集,全部使用自然语言编写。你还可以在给定相同用户的情况下的情况下两个不同模型的输出。以下是此操作的示例:

图 2.5 – 两个模型的比较示例。
对于每个模型,你还可以调整一些可以配置的参数。列表如下:
-
温度 (Temperature)(范围从 0 到 1):这控制模型响应的随机性。低温度会让你的模型更具确定性,这意味着它倾向于对相同的问题给出相同的输出。例如,如果我在温度设置为 0 的情况下多次询问我的模型OpenAI 是什么?,它将始终给出相同的答案。另一方面,如果我对一个温度设置为 1 的模型进行相同的操作,它会尝试在措辞和风格上每次都修改其答案。
-
最大标记数 (Max tokens):这控制模型对用户提示(prompt)响应的长度(以 token 为单位)。
-
停止序列 (Stop sequences)(用户输入):这使得响应在指定的点结束,例如句子或列表的末尾。
-
高概率 (Top probabilities)(范围从 0 到 1):这控制模型在生成响应时考虑哪些 token。将其设置为 0.9 将考虑所有可能的 token 中概率最高的前 90% 的部分。有人可能会问为什么不将高概率设为 1 以选择所有最可能的 token 呢?答案是,当模型置信度较低时,即使在得分最高的 token 中,用户可能仍然希望保持多样性。
-
频率惩罚 (Frequency penalty)(范围从 0 到 1):这控制生成响应中相同 token 的重复次数。惩罚越高,在同一个响应中看到相同 token 出现一次以上的概率就越低。惩罚根据 token 到此为止在文中出现的频率比例减少概率(这与下一个参数的关键区别)。
-
存在惩罚 (Presence penalty)(范围从 0 到 2):这与上一个参数类似但更严格。它减少了重复到此为止在文中出现过的任何 token 的概率。因为它比频率惩罚更严格,存在惩罚也增加了在响应中引入新话题的可能性。
除了在 Playground 中尝试 OpenAI 模型之外,你始终可以在自定义代码中调用模型 API 并模型嵌入到你的应用程序中。事实上,在 Playground 的右角,你可以点击 View code(查看代码)并按所示导出配置:
助手 (Assistants)
OpenAI 助手可以看作是一种更快速、更方便开发 AI 代理(在第 1 章中介绍)的方法。事实上,助手可以定义为由 LLM 驱动的实体,拥有一套需要遵循的指令和一套可使用的工具或插件——基本上与 AI 代理的定义相同!
在 OpenAI 助手的情况下,它们带有三个预构建工具:
-
文件搜索 (File Search):它允许用户上传自定义文档,以便助手可以导航这些文档以完成查询。它基于 RAG 框架运行。
-
函数调用 (Function Calling):它允许用户定义一组自定义函数,助手可以调用这些函数来完成给定任务。
-
代码解释器 (Code Interpreter):指的是助手针对提供的文档运行代码的能力(例如需要数学计算的电子表格或分析论文),或简单解决用户提供的复杂任务(例如复杂的数学问题)。
在接下面的图中,你可以看到一个名为“Chat with PDF”的助手的示例,它专门根据提供的文档进行回答(在我的情况下,我上传了 Hugo Touvron 等人的论文“LLaMA: Open and Efficient Foundation Language Models”)。

图 2.8 – OpenAI 助手的示例。
从前面的屏幕截可以看出,助手能够从提供的文档中检索知识来回答我的问题。事实上,我的问题相当模糊,因为“毒性”(toxicity)一词可以指多个领域;尽管如此,助手知道将提供的文档作为主要信息来源。
补全 (Completions)
此部分指的是类被称为“基础模型”(base models)的模型,如 GPT-3,它们是构建所谓的“助手模型”(或我们之前看到的聊天模型)的基础。例如,聊天模型 GPT-3.5-turbo(ChatGPT 背后的模型)是基础模型 GPT-3 的微调版本。
补全(基础)模型设计用于对提示生成单个响应,这使得它们适用于文本生成和摘要等任务,而无需在多次交互中维护上下文。相比之下,聊天(助手)模型针对交互式对话进行了优化,能够在多轮会话中维护上下文,是聊天机器人和虚拟助手等应用的理想选择。
在下面你可以看到 Playground 中一个典型补全任务的示例:

图 1: Playground 中的补全任务示例
如你所见,当我输入“今天我我去了一家杂货店并”,模型用最可能的词补全了句子。今天,补全模型很少使用,因为它们的性能被聊天模型超越,但它们可以进一步微调以适应特定的用例(我们稍后会介绍微调)。
- 文本转语音 (Text to speech):除了上述语音转文本模型
Whisper外,OpenAI 还发布了其TTS(文本转语音)模型,可以在 Playground 中直接测试。
让我们来看一个示例:

图 2: 在 Playground 中使用 OpenAI TTS 模型的示例
如上面的图片所示,你可以选择生成音频的语音、模型、速度和格式。
所有之前的模型都是预构建的,这意味着因为它们已经在海量知识库上进行了预训练。
然而,你有一些方法可以让你的模型更加定制化并适应你的用例。
第一种方法嵌入在模型的设计中,它涉及通过少样本学习方法(few-learning approach)为模型提供上下文(我们将在书中后面重点介绍这种技术)。具体来说,你可以要求模型生成一篇与你已经写的文章模板和词汇相类似的文章。为此,你可以为模型提供生成文章的查询以及之前的文章作为参考或上下文,以便模型为你的请求做好更好的准备。
以下是它的示例:

图 2.11 – OpenAI Playground 中带有少样本学习的对话示例
在之前的示例中,我指示模型仅输出推文情感的标签,并为它提供了三个示例来演示如何操作。
第二种方法更为复杂,被称为 fine-tuning(微调)。微调是将预训练模型适配到新任务的过程。
在微调过程中,预训练模型的参数会发生变化,这种变化通过调整现有参数或添加新参数,以更好地拟合新任务的数据。这是通过在特定于新任务的较小标注数据集上训练模型来实现的。微调背的核心思想是利用从预训练模型中学到的知识,并将其微调到新任务中,而不是从零开始训练模型。请看下图:

图 2.12 – 模型微调
在前面的图中,你可以看到一个关于微调如何在 OpenAI 预构建模型上工作的原理图。其核心是:你拥有一个具有通用权重或参数的预训练模型。然后,你向模型输入自定义数据,通常是以此处所示的 key-value 提示词(prompts)和完成项(completions)的形式:
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
...
一旦训练完成,你将获得一个自定义模型,它在特定任务上表现特别出色,例如对你公司的文档进行分类。
微调的妙处在于,你可以为自己的用例定制预构建模型,而无需从零开始重新训练,同时可以使用较小的训练数据集,从而缩短训练时间并减少计算量。同时,模型保留了通过原始训练(即在大规模数据集上进行的训练)学到的生成能力和准确性。
在本节中,我们概述了 OpenAI 向公众提供的模型,从你可以在 Playground 中直接尝试的(GPT、Codex)到更复杂的模型(如嵌入模型 embeddings)。我们还了解到,除了使用预构建状态的模型外,你还可以通过微调对它们进行定制,并提供一组示例供学习。
在接下来的章节中,我们将专注于这些惊模型的背景,从它们背后的数学原理开始,一直到让 ChatGPT 可能的伟大发现。
ChatGPT:顶尖技术
2022 年 11 月,OpenAI 宣布了其对话 AI 系统 ChatGPT 的网页预览版并向公众开放。这引发了来自专家、机构和普通用户的巨大热潮——在发布仅 5 天后,该服务的用户量就达到了 100 万!
在开始介绍 ChatGPT 之前,我将让它自我介绍一下,这是发布几天后的快照:

图 2.25 – ChatGPT 在 2022 年 11 月自我介绍。
如上所述,ChatGPT 的第一个版本构建在一个高级语言模型之上,该模型使用了 GPT-3 的修改版本,并专门针对对话进行了微调。这个微调后的版本被称为 GPT-3.5-turbo。优化过程涉及“基于人类反馈的强化学习”(Reinforcement Learning with Human Feedback, RLHF),这是一种利用人类输入来训练模型展现理想对话行为的技术。
我们将 RLHF 定义为一种机器学习方法,算法通过接收人类的反馈来学习执行某任务。算法被训练用于做出决策,以最大化人类提供的奖励信号,而人类则提供额外的反馈以改进算法的性能。当任务复杂到传统编程无法处理,或者预期结果难以预定义时,这种方法非常有用。
这里相关的区别在于,ChatGPT 在训练过程中包含了人类参与(humans in the loop),以便使其与用户保持一致(aligned)。通过引入 RLHF,ChatGPT 的设计旨在以更自然、更有互动的方式理解并响应人类语言。
同样的
RLHF机制也被用于我们可以认为是 ChatGPT 前身的模型——InstructGPT。在 OpenAI 研究人员于 2022 年 1 月发表的相关论文中,InstructGPT被介绍为一类在遵循英语指令方面优于GPT-3的模型。
ChatGPT 直如今仍作为任何人都可以使用的免费应用程序可用;但自 2023 年 2 月起,OpenAI 宣布了一个名为 ChatGPT Plus 的新付费版本,每月费用 20 美元,为订阅者提供了若干优势,包括访问最新模型、最快的响应时间、一套强大的插件集,以及在 GPTs 游乐场中创建自己的助手的可能性。
让我们快速浏览一下当前日期的 ChatGPT 用户界面:

图 3: chatgpt.com 的 ChatGPT 首页
让我们双击每个部分:
-
你可以决定
ChatGPT后端使用的模型。在我情况下,我设置了GPT-4o,它仅对付费订阅开放。这是我们在整本书中将使用的模型。 -
系统向用户推荐一组预构建的提示词,帮助他们熟悉应用程序。
-
文本框是用户提问的地方。注意左上角有一个小回针图标:它表示模型可以上传的文件。这些文件可以从本地上传,也可以从
Google Drive或OneDrive等云存储检索。 -
在左侧边栏中,你可以看到与
ChatGPT的聊天列表。这是一个非常有用的工具,因为在每个聊天中,通过你与模型的多次交互,创建了模型感知的上下文。这意味着如果你想继续之前的对话,只需打开相应的聊天即可,无需再次描述整个场景。 -
最近添加的一个很棒的功能是与其他用户共享工作空间的功能,以便你与团队成员协作。请注意这是一个额外的付费功能,不包含在 Plus 计划中,每月额外收费 5 美元。
-
最后,
ChatGPT Plus提供了创建 GPTs 的可能性,这些是你为特定功能定制的个性化助手。你可以决定将 GPT 设置为私有,或将其发布到 GPTs 商店供任何人使用和评分。我们将在本书后后的专门章节中介绍 GPTs。
最后,ChatGPT Plus 带有两个可以根据需要自动调用的内置插件:网络搜索插件(用于检索最新信息)和代码解释器插件(用于处理分析文档或执行复杂的数学任务)。
在本书中,我们将利用 ChatGPT Plus 展示最新模型和功能的特性,尽管如此,我们涵盖的大多数示例通过免费版的 ChatGPT 实现(该版本目前由 GPT-3.5-turbo 驱动)。
ChatGPT 架构和训练方法的持续发展和改进有望进一步推向语言处理的边界。
总结
在本章中,我们回顾了 OpenAI 的历史、研究领域以及最新进展,直到 ChatGPT。我们深入探讨了用于测试的 OpenAI Playground 以及我们如何利用 OpenAI 的所有模型系列。
通过对 OpenAI 模型的初步了解,我们将看到测试或将预训练模型嵌入到你的应用程序中是多么简单:这里的关键因素在于,你不需要强大的硬件和花费数小时的时间来训练模型,因为它们已经可以供你使用,并在需要时还可以通过少量示例进行自定义。
在下一章中,我们将开始介绍本书的Part 2,在 ways 我们将看到 ChatGPT 在各个领域的实际应用,以及如何释放它的潜力。你将学习如何通过正确设计提示(prompts)从 ChatGPT 中获得最高价值,如何提高你的日常工作效率,以及它如何成为开发者、营销人员和研究人员优秀的项目助手。
参考文献
-
Radford, A., & Narasimhan, K. (2018). Improving language understanding by generative pre-training.
-
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. ArXiv.
doi.org/10.48550/ arXiv.1706.03762OpenAI. Fine-Tuning Guide. OpenAI platform documentation.platform.openai.com/docs/guides/fine-tuning。
3 理解提示词设计
在 Discord 上加入我们的书籍社区

在之前的章节中,我们在提到 ChatGPT 以及通用的 LLM 模型的用户输入时,多次使用了 prompt(提示词)一词。
由于提示词对 LLM 的性能有巨大影响,因此提示词工程(prompt engineering)是设计 LLM 驱动的应用时项至关重要的活动。事实上,有多种技术可以实现,不仅可以优化你的 LLM 响应,还可以减少与幻觉(hallucination)和偏差相关的风险。
在本章中,我们将涵盖提示词工程领域的新兴技术,从基础方法到高级框架。具体来说,我们将介绍以下主题:
-
提示词工程简介
-
零样本、单样本和少样本学习(Zero, one and few shot learning)
-
提示词工程的基本原理
-
提示词工程的高级技术
-
通过提示词工程缓解风险和幻觉
-
处理提示注入(prompt injections)
在本章结束时,你将拥有为你的 LLM 驱动的应用构建功能且稳健的提示词的基础,这在接下来的章节中也会相关。
什么是提示词工程?
在解释什么是提示词工程之前,让我们从提示词(prompt)的原子定义开始。
提示词是一种文本输入,引导 LLM 的行为以生成文本输出。在 LLM 和 LLM 驱动的应用的语境下,我们可以将提示词分为两类:
-
用户在与 LLM 交互时输入的提示词。例如,一个提示词可能是“给我详细解释一下质子”,或者“生成一个运行马拉松的跑步训练计划”。下面你可以看到一个简单的用户提示词示例:
![]()
![图 1:用户提示词示例。]()
图 1:用户提示词示例。
你会听到人们简单地将此组件称为“prompt”、“query”(查询)或“user’s input”(用户输入)。
-
无论用户的查询如何,指示模型以某种方式运行的提示词。这指的是我们提供给模型的自然语言指令集,使其在与终端用户交互时以某种方式运行。你可以将其理解为 LLM 的某种“后端”,即由应用程序开发者处理而非最终用户处理的内容。
在接下来的图片中,你可以看到一个示例,我们向模型提出了上述问题,但这次我们添加了一个系统消息(system message)来指示模型以某种方式运行:



图 2:元提示词示例。
我们将这类提示词称为“metaprompt”(元提示词)或“system message”(系统消息)。
从此以后,提示词工程是指设计有效提示词的过程,旨在从 LLM 中引导高质量且相关的输出。提示词工程需要创造力、对 LLM 的理解以及精确性。

图 3:通过提示词工程使 LLM 专业化的示例。
在过去的几年里,提示词工程本身变成了一门全新的学科,这证明了与这些模型交互需要一套以前不存在的新技能和能力。“提示词艺术”(art of prompting)已成为在企业场景下构建生成式 AI 应用的顶级技能;然而,对于在日常任务中使用 ChatGPT 或类似 AI 助手的个人用户来说,它也非常有用,因为它极大提高了结果的质量和准确性。
在接下来的章节中,我们将看到一些如何为你的 LLM 应用构建高效稳健的提示词和元提示词的示例。请注意,我们将涵盖的所有技术都可以嵌入到用户提示词和元提示词中,取决于你的具体需求。为了做到这一点,我将利用 OpenAI Playground,以便能够更灵活地在两个层级的提示上进行操作。
零样本、单样本和少样本学习——Transformer 的典型
在之前的章节中,我们提到了 LLM 通常以预训练格式出现。它们已经在海量数据上进行了训练。
然而,这并不意味着它们现在不能学习了。在第 2 章中,我们看到自定义 OpenAI 模型并使其更适用于特定任务的方法是微调(fine-tuning)。
微调是将预训练模型适配到新任务的过程。在微调中,预训练模型的参数会被更改,即通过调整现有参数或添加新参数来适应数据。这是通过在特定任务的较小数据集上训练模型来实现的。微调的核心思想是利用从预训练模型中学到的知识并使其新任务,而不是从头训练模型。
微调是一个正式的训练过程,需要训练数据集、计算能力和一些时间(取决于数据量和计算实例)。
这就是为什么测试另一种让模型在特定任务中变得更技能的方法:样本学习(shot learning):其思想是让模型从示例学习,而不是整个数据集。这些示例是我们希望模型响应的方式,以便模型不仅学习内容,还学习在其响应中使用的格式、风格和分类。
此外,样本学习直接通过提示词发生(正如我们在后续场景中看到的),因此整个过程更不耗时且更容易执行。
提供的示例数量决定了我们所指的样本学习水平。换句话说,如果不提供示例则为零样本(zero-shot),提供一个示例则为单样本(one-shot),提供 2-3 个以上的示例则为少样本(few-shot)。
让我们关注每一个场景:
-
零样本学习。在这种类型的学习中,模型被要求执行一项它未见过任何训练示例的任务。模型必须依靠先验知识或关于任务的通用信息来完成任务。例如,零样本学习的方法可以是要求模型生成描述,正如我的提示词中所定义的:
![图 4.3 零样本学习示例]()
图 4.3 零样本学习示例
- 单样本学习(One-shot learning):在这种学习类型中,模型对于每个要求执行的新任务只获得一个示例。模型必须利用其先验知识从这个单一示例中泛化,以完成任务。如果我们考虑之前的示例,我可以在要求模型生成新内容之前,为它提供一个
prompt-completion示例:

图 4.4 – 单样本学习示例
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
注意,我提供示例的方法与用于微调的结构相似:
- 少样本学习(Few-shot learning):在这种学习类型中,模型对于每个要求执行的新任务会被给出一个少量示例(通常在 2 到 5 个之间)。模型必须利用其先验知识从这些示例中泛化以完成任务。让我们继续之前的示例,为模型提供更多示例:

图 4.5 – 包含三个示例的少样本学习示例
少样本学习的优点在于,你还可以控制模型输出的呈现方式。你也可以为模型提供一个你希望输出格式的模板。例如,考虑以下推文分类器:

图 4.6 – 用于推文分类器的少样本学习。这是来自 https://learn.microsoft.com/en-us/azure/cognitive-services/openai/how-to/completions 原始脚本的修改版本。
让我们检查上面的图。首先,我为 ChatGPT 提供了一些带标签的推文示例。然后,我提供了相同的推文,但使用了不同的数据格式(列表格式),以及同样格式的标签。最后,我以列表格式提供了未标记的推文,以便模型以列表形式返回标签。
标签。
少样本学习(Shot-learning)的可能性是无限的——这取决于测试以及在寻找合适的提示词(prompt)设计时的一点耐心。
如之前提到的,重要的是记住这些学习形式与传统的监督学习以及微调是不同的。在少样本学习中,目标是让模型能够从极少数示例中学习,并从这些示例泛化到新任务。
现在我们已经学会了如何让 OpenAI 模型从示例中学习,让我们关注如何正确定义我们的提示词,以使模型的响应尽可能准确。
提示词工程原则
传统上,在计算和数据处理的背景下,我们经常使用“垃圾进,垃圾出(garbage in, garbage out)”这一说法,意味着输出的质量取决于输入的质量。如果向系统输入了错误或低质量的数据(垃圾),输出也将是有缺陷或毫无意义的(垃圾)。
在涉及提示词(prompting)时,情况也类似:如果我们希望从 LLM 获得准确且相关的结果,我们需要提供高质量的输入。然而,构建良好的提示不仅取决于响应的质量。事实上,我们可以构建良好的提示词来:
-
最大化 LLM 响应的相关性
-
指定响应的类型、格式和风格
-
提供对话上下文
-
减少内部偏见并提高公平性和包容性
-
减少幻觉
在大语言模型(LLMs)的语境下,“幻觉(hallucination)”指的是生成事实上错误、无意义或未基于训练数据的文本或响应。当 LLM 产生听起来自信但错误或虚假的信息时,就会发生幻觉。幻觉的产生由于这些模型的概率性质,它们可能会根据模式而非验证的事实来预测下一个词或短语。
让我们看看实现这些结果的基本技术。
清晰的指令
给出清晰指令的原则是为模型提供足够的信息和指导,使其正确且高效地执行任务。清晰的指令应该包含以下元素:
-
任务的目标或目的,例如“写一首诗”或“总结文章”。
-
预期输出的格式或结构,例如“使用带有押韵字的四行诗”或“使用项目符号,每项不超过 10 个字”。
-
任务的约束或限制,例如“不要使用任何脏话”或“不要从源文件中复制任何文本”。
-
任务的上下文或背景,例如“这首诗是关于秋天的”或“这篇文章来自某科学期刊”。
假设,我们希望模型从文本中提取任何类型的指令,并以项目列表的形式向我们返回教程。此外,如果提供的文本没有指令,模型应该告知我们。
为了尝试这样做,让我们利用 OpenAI Playground 中的聊天部分,这样我们就可以同时提供元提示词(metaprompt)和提示词(prompt)。
对于这种场景,我将设置以下元提示词:
你是一个 AI 助手,通过根据文本生成教程来帮助人类。
你将提供一段文本。如果文本包含关于如何进行某事的任何指令,请生成一个列表格式的教程。
否则,告知用户文本不包含任何指令。
这是我们的用户提示词:
为了准备来自意大利热那亚的已知酱汁,你可以从烤松子开始,然后在厨房研钵中与罗勒和大蒜一起粗略切碎。然后,在厨房研钵中加入一半油,并用盐和胡椒调味。
最后,将青酱转移到碗中,拌入碎的帕玛森奶酪。
让我们看看它是如何工作的:

图 4: 元提示词中清晰指令的示例。
注意,如果我们给模型传递另一个不包含任何指令的文本,它将按照我们指示的方式做出回答:

图 5: 聊天模型遵循指令的示例。
通过给出清晰的指令,你可以帮助模型理解你想让它做什么以及你想让怎么做。这可以提高模型输出的质量和相关性,并减少进一步修订或修正的需求。
然而,有时仅仅有清晰度是不够的。我们可能需要推断 LLM 的思维方式,使其在执行任务时更加健壮。在下一节中,我们将检查这种技术之一,它在完成复杂任务时非常有用。
将复杂任务拆分为子任务
提示词工程是一项设计有效输入以供大语言模型(LLMs)执行各种任务的技术。有时,任务对于单个提示词来说过于复杂或模糊,此时将它们拆分为由不同提示词解决的更简单的子任务更好。
以下是将复杂任务拆分为子任务的一些示例:
-
文本摘要:一项涉及为长文本生成简洁准确摘要的复杂任务。该任务可以拆分为若干子任务,例如:
-
从文本中提取要点或关键词。
-
以连贯流畅的方式重写要点或关键词。
-
修剪摘要以符合所需的长度或格式。
-
-
机器翻译:一项涉及将文本从一种语言翻译成另一种语言的复杂任务。该任务可以被拆分为若干子任务,例如:
- 检测文本的源语言。
-
将文本转换为中间表示,并保留原始文本的含义和结构。
-
从中间表示生成目标语言的文本。
-
诗歌生成:一项创意任务,涉及创作遵循特定风格、主题或情感的诗歌。该任务可以分为若干子任务,例如:
-
为诗歌选择诗歌形式(如十四行诗、俳句、五味打诗等)和韵律(如 ABAB、AABB、ABCB 等)。
-
根据用户的输入或偏好为诗歌生成标题和主题。
-
生成符合所选形式、韵律和主题的诗行或诗节。
-
对诗歌进行精炼和润色,以确保连贯性、流畅性和原创。
-
-
代码生成:一项技术任务,涉及生成执行特定功能或任务的代码段。该任务可以分为若干子任务,例如:
-
为代码选择编程语言(如
Python、Java、C++等)以及框架或库(如TensorFlow、PyTorch、React等)。 -
根据用户的输入或规范为代码生成函数名以及参数列表和返回值。
-
生成实现代码逻辑和功能的函数体。
-
添加注释和文档以解释代码及其用法。
-
-
让我们考虑以下示例。我们将为模型提供的一篇短文章,并要求它按照这些指令进行总结:
You are an AI assistant that summarize articles.
To complete this task, do the following subtasks:
Read the provided article context comprehensively and identified the main topic and key points
Generated a paragraph summary of the current article context that captures the essential information and conveys the main idea
Print each step of the process.
这是我将提供的短文:
大语言模型(LLMs)作为人工智能的一个子集,通过展现出前所未有的理解和生成类人文本的能力,彻底改变了自然语言处理领域。这些模型在包含多样化语言输入的海量数据集上进行训练,使其能够针对广泛的主题生成连贯且上下文相关的响应。通过利用 transformers 等架构,GPT-3 及其后续者等 LLM 可以完成文本补、回答问题、执行翻译,甚至参与复杂的对话。它们的应用范围从自动化客户支持、内容创作到先进的研究和教育工具。尽管拥有惊人的能力,LLM 也面临着挑战,包括训练数据中固有的偏差倾向以及生成误导性或虚假信息的风险。随着 LLM 的不断进步,在伦理 AI 研究和部署策略方面的持续的努力对于负责任且有效地利用其益处至关重要。
让我们看看模型是如何工作的:

图 6:OpenAI GPT-4o 将任务拆分为子任务以生成摘要的示例。
将复杂任务拆分为更简单的子任务是一种强大的技术,尽管如此它并没有解决 LLM 生成内容的主要风险之一,即输出错误结果。在接下来的两个章节中,我们将看到一些旨在解决这一风险的技术。
请求理由(Ask for justification)
LLM 的构建方式是它们根据之前的标记(token)预测下一个标记,而不回溯其已生成的内容。这可能导致模型向用户输出错误内容,但方式非常具有说服力。如果 LLM 驱动的应用没有为该响应提供特定的引用,可能很难验证其背后的事实。
因此,在提示词(prompt)中指定通过一些反思和理由来支持 LLM 的答案,可能会提示模型从其行为中恢复。
此外,请求理由在答案正确的情况下也可能有用,只是因为我们并不知道 LLM 背后的思考过程。例如,假设我们想让 LLM 解决谜语。为了做到,我们可以这样指示它:
You are an AI assistant specialized in solving riddles.
Given a riddle, solve it the best you can.
Provide a clear justification of your answer and the reasoning behind it.
如你所见,我在元提示词(metaprompt)中指定 LLM 为其答案提供理由,并将其推理社交化。让我们用下面的谜语来测试一下:
什么有脸有只手,但没有手脚?
输出:

图 7:OpenAI 的 GPT-4o 在解决谜语后提供理由的示例。
理由是使你的模型更可靠、更健壮的佳工具,因为它们“迫使”模型重新思考其输出,同时为我们提供了推理是如何设定以解决问题的视角。
通过类似的方法,我们也可以在不同的提示词层面进行干预,以提高 LLM 的性能。例如,我们可能会发现模型正在系统地以错误的方式处理数学问题,因此我们可能想在元提示词层面上直接建议正确的方法。另一个例子可能是要求模型生成多个输出——以及它们对应的理由——以评估不同的推理技术,并在元提示词中选择最佳的一个。
在下一节中,我们将关注这些示例之一,更具体说是生成多个输出然后选择最可能的一个的可能。
生成多个输出,然后使用模型选择最佳的一个
正如我们在上一节中看到的,LLM 的构建方式是它们根据之前的标记预测下一个标记,而不回溯其已生成的内容。如果是这种情况下,如果采样的标记是错误的(换句话说,如果模型运气不好),LLM 将继续生成错误的标记,从而产生错误的内容。现在坏消息是,与人类不同,LLM 无法从错误中自动恢复。这意味着,如果我们询问它们,它们会承认错误,但我们需要明确提示它们进行思考。
克服这一局限性的方法是扩大选择正确标记的概率空间。与其只生成一个响应,我们可以提示模型生成多个响应,然后选择最适合用户查询的一个。这为我们的 LLM 分配两个子任务:
-
为用户查询生成多个响应;
-
比较这些响应并根据我们可以在元提示词中指定的标准选择最佳的一个。
让我们来看一个示例,接上上一节检查过的谜语:
You are an AI assistant specialized in solving riddles.
Given a riddle, you have to generate three answers to the riddle.
For each answer, be specific about the reasoning you made.
Then, among the three answer, select the one which is most plausible given the riddle.
在这种情况下,我提示模型为谜语生成三个答案,然后给出最可能的那个,并说明理由。让我们看看结果:

图 8:GPT-4o 生成三个可能的答案并选择最可能的一个、提供理由的示例。
如前所述,强制模型用不同的方法解决问题是收集多个推理样本的一种方法,这些样本可以作为元提示词(metaprompt)中的进一步指令。例如,如果我们希望模型总是对问题提出不是最直观的解决方案——换句话说,如果我们希望它能“以不同方式思考”——我们可能会强制它以 N 种方式解决问题,然后将最具创意的推理作为元提示词中的框架。
我们将检查的最后一个元素是我们想要赋予元提示词的整体结构。事实上,在之前的示例中,我们看到了一个包含某些陈述和指令的示例系统消息。在下一节中,我们将看到这些陈述和指令的顺序和“强度”并不是不变的。
使用分隔符
涵盖的最后一个原则与我们想要赋予元提示词的格式有关。这有助于我们的大语言模型(LLM)更好地理解其意图,并建立章节与段落之间的关系。
为了实现这一点,我们可以在提示词中使用分隔符。分隔符可以是任何字符或符号序列,它能清晰地映射到一个模式(schema)而不是一个概念。例如,我们可以考虑以下序列作为分隔符:
-
>>>> -
==== -
------ -
#### -
`
等等。让我们考虑一个元提示词,其目标是指示模型将用户的任务转换为 Python 代码,并提供一个操作示例。
You are a Python expert that produces python code as per user's request.
===>START EXAMPLE
---User Query---
给我一个打印文本字符串的函数。
---User Output---
下面你可以找到描述的函数:
```def my_print(text):
#returning the printed text
return print(text)
<===END EXAMPLE
让我们看看它是如何工作的:

图 9:在系统消息中使用分隔符的模型示例输出。
如你所见的,它还像系统消息中显示的那样用反引号打印了代码。
到目前为止的所有原则都是通用规则,可以使你的基于 LLM 的应用程序更加健壮。在接下来的章节中,我们将看到一些高级的提示工程技术,这些技术在将答案提供给最终用户之前,解决了模型推理和思考答案的方式。
## 高级技术
在之前的章节中,我们涵盖了一些提示工程的基础技术。无论你开发的是什么类型的应用程序,都应该记住这些技术,因为它们是提高你 LLM 性能的通用最佳实践。
另一方面,还有一些针对特定场景实施的高级技术,我们将在接下来的章节中介绍。
### 思维链 (Chain of Thoughts)
由 Wei 等人在论文《Chain-of-Thought Prompting Elicits Reasoning in Large Language Models》中引入的思维链(CoT)是一种通过中间推理步骤实现复杂推理能力的技术。它还鼓励模型解释其推理过程,“强制”它不要太草率,从而避免给出错误回答的风险(正如我们在之前的章节中看到的)。
假设我们想提示 LLM 求解一元方程。为了做到,我们将为它提供一个通用的推理列表作为元提示词:
要求解通用的一元方程,请遵循以下步骤:
1. 识别方程: 首先识别你想要求解的方程。它应该是 "ax + b = c" 的形式,其中 'a' 是变量的系数,'x' 是变量,'b' 是常数,'c' 是另一个常数。
2. 隔离变量: 你的目标是在方程的一侧隔离变量 'x'。为了做到,请执行以下步骤:
a. 加减常数: 在方程两边加上或减去 'b',将常数移到一侧。
b. 除以系数: 将方程两边除以 'a' 以隔离 'x'。如果 'a' 为零,方程可能没有唯一解。
3. 简化: 尽可能地简化方程的两边。
4. 求解 'x': 一旦 'x' 在一侧被隔离,你就得到了解。它将是 'x = 值' 的形式。
5. 检查你的解: 将找到的 'x' 值代回原始方程,确保它满足方程。如果是,你就找到了正确解。
6. 表示解: 以清晰简洁的形式写下解。
7. 考虑特殊情况: 注意可能无解或无限解的特殊情况,特别是当 'a' 等于零时。
方程:
让我们看看它是如何工作的:

图 10:模型使用 CoT 方法求解方程的输出。
注意,你也可以将其与少样本提示(few-shot prompting)结合,以便在回答之前需要推理的更复杂任务上获得更好的结果。
通过 CoT,我们正在提示模型生成中间推理步骤。这也是我们在下一节中将检查的另一种推理技术的组成部分。
### ReAct
由 Yao 等人在论文《ReAct: Synergizing Reasoning and Acting in Language Models》中引入的 ReAct(推理与行动)是一种将推理与大语言模型相结合的通用范式。ReAct 提示语言模型生成语言推理轨迹,并接收来自网络等外部来源的观察。这允许语言模型执行动态推理,并根据外部信息快速调整计划。例如,你可以提示语言模型通过对问题进行推理,然后执行向网络发送查询,然后接收搜索结果的观察,并继续这种“思考-行动”循环,直到得出结论。
思维链和 ReAct 方法的区别在于,思维链提示语言模型为任务生成中间推理步骤,而 ReAct 提示语言模型为任务生成中间推理步骤、行动和观察。
在下面的示例中,我们不使用工具,而是将要求模型执行的任何任务称为“行动”。
这个 ReAct 元提示词可能像这样:
尽力回答以下问题。
使用以下格式:
Question: 你必须回答的输入问题
Thought: 你应该始终思考要做什么
Action: 要采取的行动
Action Input: 动作的输入
Observation: 行动的结果
...(此 Thought/Action/Action Input/Observation 可以重复 N 次)
Thought: 我现在知道了最终答案
Final Answer: 原始输入问题的最终答案
让我们看看它在简单的用户查询下是如何工作的:

图 11:ReAct 元提示词的示例
如你所见,在这种场景下,模型跳过了 Action Input 和随后的观察,因为我们没有为它提供任何可使用的工具。
这是一个很好的例子,说明了提示模型逐步思考并显式说明推理每一步如何让它在回答之前变得更加“明智”和谨慎。这也是防止幻觉的一种很好的技术。
总的来说,提示工程是一门强大的学科,尽管仍处于新兴阶段,但已经在基于 LLM 的应用程序中被广泛采用。在接下来的章节中,我们将看到这些技术的具体应用。
# 在 ChatGPT 中避免隐藏偏见风险并考虑伦理考量
ChatGPT 提供了 `Moderator API`,使其使其不会参与可能不安全的对话。`Moderator API` 是一个基于 GPT 模型的分类模型,涵盖了以下几类:暴力、自残、仇恨、骚扰和性。为此,OpenAI 使用了匿名数据和合成数据(以零样本形式)来创建合成数据。
`Moderation API` 基于 OpenAI API 中提供的更复杂的内容过滤模型版本。我们在*第 1 章*中讨论过该模型,当时看到它对假误报的处理非常保守,而不是对漏报如此保守。
然而,存在一种我们可以称之为`隐藏偏见`的现象,它直接源于模型训练的知识库。例如,关于 `GPT-3` 的主要训练数据块 `Common Crawl`,专家们认为它主要由西方国家的白男性编写。如果确实如此,我们已经面临着模型的隐藏偏见,这将不可避免地模仿一个有限且缺乏代表性的人类类别。
在《Languages Models are Few-Shots Learners》论文中,OpenAI 的研究人员 Tom Brown 等人([`arxiv.org/pdf/2005.1416`](https://arxiv.org/pdf/2005.1416))创建了一个实验设置,以调查 `GPT-3` 中的种族偏见。模型被提示包含种族类别的短语,并为每个类别生成了 800 个样本。使用 `Senti WordNet` 根据词汇共现衡量了生成文本的情感,评分范围从 -100 到 100(正分表示正性词汇,反之亦然)。
结果显示,与每个种族类别相关的情感在不同模型之间存在差异,亚洲人始终具有高情感,而黑人始终具有低情感。作者警告说,结果反映了实验设置,社会历史因素可能会影响与不同人口统计相关的情感。该研究强调了对情感、实体和输入数据之间关系进行更复杂分析的性:

图 4.14 – 跨模型的种族情感
这种隐藏偏见可能会产生不符合负责任 AI 原则的有害响应。
然而,值得注意的是,ChatGPT 以及所有 OpenAI 模型都在不断改进中。这与 OpenAI 的`AI 对齐`([`openai.com/alignment/`]( https://openai.com/alignment/ ))是一致的,其研究重点是训练 AI 系统变得有帮助、诚实且安全。
例如,如果我们要求 ChatGPT 根据人的性别和种族进行猜测,它不会满足这一具体请求,而是提供我们假设性的功能以及巨大的免责声明:

图 4.15 – GPT-4o 随时间改进的示例,因为它提供了无偏见的响应
总而言,尽管伦理原则领域在不断改进,但在使用 ChatGPT 时,我们应该始终确保输出符合这些原则且没有偏见。
ChatGPT 和 OpenAI 模型中的偏见和伦理概念在负责任 AI 的整个主题中有更广泛的关联,我们将在本书的最后一章关注这一点。
## 总结
在本章中,我们深入探讨了提示设计(prompt design)和工程的概念,因为这是控制 ChatGPT 以及通用大语言模型(LLMs)输出的最强大方式。我们学习了如何利用不同层次的样本学习(shot learning)来使 LLMs 更符合我们的目标。
我们从介绍提示工程的概念及其重要性开始,然后转向基本原则——包括清晰的指令、要求提供理由等。
然后,我们转向了更高级的技术,这些技术旨在塑造我们 LLM 的推理方法:少样本学习、`CoT` 和 `ReAct`。
提示工程是一门新兴学科,它正在为注入 LLM 的新类别应用程序铺平道路。在接下来的章节中,我们将看到这些技术如何在构建 LLMs 的真实应用中投入。
从下一章开始,我们将深入探讨 ChatGPT 可以提高生产力并对我们今天的工作方式产生破坏性影响的不同领域。
## 参考文献
以下是本章的参考文献:
* [`arxiv.org/abs/2005.14165`](https://arxiv.org/abs/2005.14165)
* [`dl.acm.org/doi/10.1145/3442188.3445922`](https://dl.acm.org/doi/10.1145/3442188.3445922)
* [`openai.com/alignment/`](https://openai.com/alignment/)
* [`twitter.com/spiantado/status/1599462375887114240?ref`](https://twitter.com/spiantado/status/1599462375887114240?ref)
* ReAct approach. [`arxiv.org/abs/2210.03629`](https://arxiv.org/abs/2210.03629)
* Chain of Thoughts approach. [`arxiv.org/abs/2201.11903`](https://arxiv.org/abs/2201.11903)
ChatGPT 也可以成为忠诚且自律的学习伙伴。例如,它可以帮助你总结长篇论文,以便你对所讨论的话题有一个初步的了解,或者帮助你准备考试。
具体来说,假设你正在使用由 Lorenzo Peccati 等人编写的名为 `Mathematics for Economics and Business` 的大学教材准备数学考试。在深入研究每个章节之前,你可能想对内容和讨论的主要主题有一个概述,并确认是否需要更多的先备知识,以及最重要的一点——如果我正在准备考试的话——学习它需要多长时间。你可以这样问 ChatGPT:

图 5.5 – ChatGPT 提供大学教材的概述
当涉及特定资产或个人信息的问题时,我们增加了 ChatGPT 幻觉的风险。事实上,在在上一个例子中,模型能够回答是因为该书的摘要显然是训练集的一部分,但模型并没有该书的完整内容,因为其在互联网上无法免费获取。如果是这种情况,指定模型浏览网络来回答这些特定问题可能是一个很好的做法。让我们来看一个例子:

图 1:ChatGPT 调用网络插件的示例
如你所看到的,现在 ChatGPT 调用了网络插件并搜索了 5 个文件,还提供了它浏览过的链接。
你也可以要求 ChatGPT 针对你刚刚学习的材料向你提问题:

图 5.6 – ChatGPT 扮演教授的示例
> `Act as`(充当…………技术是高效提示技巧的绝佳示例,它可以列入 `第 3 章` 描述的示例中。
现在,让我们来看看更多使用 ChatGPT 执行特定任务的例子,包括文本生成、写作辅助和信息检索。
## 生成文本
作为一个语言模型,ChatGPT 特别适合根据用户的指令生成文本。例如,你可以要求 ChatGPT 生成针对特定受众的电子邮件、草稿或模板:

图 5.7 – ChatGPT 生成的电子邮件示例
另一个例子可能是要求 ChatGPT 为你必须准备的演示文稿创建一个演讲大纲(我在这里只放第一张幻灯片示例):

图 5.8 – ChatGPT 生成的幻灯片日程和结构
你还可以以此方式生成关于趋势话题的博客文章或文章。这是一个示例:

图 5.9 – ChatGPT 生成的带有相关标签和 SEO 关键词的博客文章示例
我们甚至可以让 ChatGPT 缩小文章的大小以适应推文。以下是操作方法:

图 5.10 – ChatGPT 将文章缩短为 Twitter 推文
最后,ChatGPT 还可以生成视频或戏剧剧本,包括舞台设计和建议的剪辑。底图显示了一个包含舞台设计和演员的剧场对话示例:

图 5.11 – ChatGPT 生成的带有舞台设计的剧场对话
我只提供了 ChatGPT 生成的四个场景中的一个,让你对结局保持悬念……
总的来说,每当需要从零生成新内容时,ChatGPT 在提供初稿方面做得非常出色,初稿可以作为进一步优化的起点。
然而,ChatGPT 还可以通过提供写作辅助和翻译来支持现有内容,我们将在下一节看到。
## 提高写作技巧和翻译
有时,与其生成新内容,你可能想重新审视现有的文本。可能是为了改进风格、改变受众、语言翻译等。
让我们来看一些例子。假设我草拟了一封邮件,邀请我的客户参加网络研讨会。我写了两个短句。在这里,我想让 ChatGPT 改进这封邮件的格式和风格,因为目标受众是高管:

图 5.12 – ChatGPT 针对高管受众修改的电子邮件示例
现在,让我们问同样的问题,但针对不同的受众:

图 5.13 – ChatGPT 生成的针对不同观众的相同邮件示例
ChatGPT 还可以对你的写作风格和结构提供反馈。
例如,假设你为标题为 `The History of Natural Language Processing` 的文章写了引言,你想就写作风格及其与标题的一致性获得反馈:

图 5.15 – ChatGPT 对文章引言提供反馈的示例
如你所看到的,ChatGPT 不仅为整篇文章提供了反馈和建议,还给我了一个整合了所有建议的修订引言。
让我们也要求 ChatGPT 对其`细化示例` 说得更具体:

图 5.16 – ChatGPT 对提到内容的详细说明
我也有兴趣知道引言与标题是否一致,或者我是否偏离了方向:

图 5.17 – ChatGPT 提供引言与标题一致性的反馈
最后一个让我印象深刻。ChatGPT 足够聪明,能够发现我的引言中没有具体提到 NLP(自然语言处理)。然而,它为稍后讨论该主题设定了预期。这意味着 ChatGPT 在文章结构方面也具有专业知识,并且在应用判断时非常精准,知道这只是一个引言。
现在让我们揭晓本章 ChatGPT 的最后项技能。事实上,ChatGPT 也是一个出色的翻译工具。它认识 95 种语言(如果你怀疑你的语言是否被支持,可以直接询问 ChatGPT)。然而,这里存在一种考虑:在已经有了 Google Translate 等尖端工具的情况下,ChatGPT 翻译的额外价值何?
为了回答这个问题,我们必须考虑一些关键区别以及如何利用 ChatGPT 的翻译功能:
* ChatGPT 可以捕捉意图。这意味着你也可以跳过翻译阶段,因为这是 ChatGPT 在后台可以完成的。例如,如果你写一个提示来生成法语社交媒体帖子,你可以用任何你喜欢的语言编写该提示——ChatGPT 会自动检测(无需提前指定)并理解你的意图:

图 5.18 – ChatGPT 生成与输入语言不同的输出示例
* ChatGPT 能够捕捉俚语或成语更深层次的含义。这使得翻译不再是字面直译,从而能够保留其深层含义。例如,让我们考虑英语表达 `It’s not my cup of tea`,它表示某物符合符合人的喜好或偏好。让 ChatGPT 和 Google Translate 将其翻译成意大利语:

图 5.19 – ChatGPT 与 Google Translate 从英语翻译成意大利语对比
如你所示,ChatGPT 可以提供几个与原句等价的意大利语成语,同样以俚语的形式呈现。相比之下,Google Translate 执行了字面翻译,忽略了成语的真实含义。
* 与任何其他任务一样,你总是可以为 ChatGPT 提供上下文。因此,如果你希望翻译具有特定的俚语或风格,你始终可以在提示词中指定。或者,更有趣的是,你可以要求 ChatGPT 用讽刺的语气翻译你的提示词:

图 5.20 – ChatGPT 用讽刺语气翻译提示词的示例。提示词原始内容取自 OpenAI 的维基百科页面:https://it.wikipedia.org/wiki/OpenAI
所有这些场景都突出了 ChatGPT 以及整个 OpenAI 模型的一个核心杀手锏功能。由于它们代表了 OpenAI 定义的 `通用人工智能` (`AGI`) 的体现,它们的设计初衷并非为了专注于单一任务(即被局限于单一任务)。相反,它们旨在动态地服务于多种场景,以便你通过一个模型处理广泛的用例。
总之,ChatGPT 不仅能够生成新文本,还能操纵现有材料以满足你的需求。它已被证明在语言之间的翻译方面非常精确,同时保持术语和特定语言表达的完整性。
在下一节中,我们将看到 ChatGPT 如何帮助我们检索信息和竞争情报。
## 快速信息检索与竞争情报
信息检索和竞争情报是 ChatGPT 另一个改变游戏规则的领域。ChatGPT 检索信息的第一个例子就是它目前最流行的使用方式:作为搜索引擎。每当我们问 ChatGPT 某些问题时,它都能从其知识库中检索信息并以原创的方式进行重构。
一个例子是要求 ChatGPT 为我们可能感兴趣阅读的书籍提供快速摘要或评论:

图 5.21 – ChatGPT 提供书籍摘要和评论的示例
或者,我们可以根据自己的偏好为想要读的新书寻求一些建议:

图 5.22 – ChatGPT 根据我的偏好推荐书单的示例
此外,如果我们在提示词中加入更具体的信息,ChatGPT 可以作为一个工具,引导我们找到研究或学习的正确参考文献。
具体来说,你可能想要快速检索关于你想要进一步学习的主题的一些背景参考文献——例如,前馈神经网络。你可能会要求 ChatGPT 指向一些介绍广泛讨论该主题的网站或论文:

图 5.23 – ChatGPT 列出相关参考文献
如你所示,ChatGPT 能够为我提供相关的参考文献以开始学习该主题。然而,在竞争情报方面,它还可以更进一步。
假设我正在写一本名为 `Introduction to Convolutional Neural Networks – an Implementation with Python` 的书。我想对市场上的潜在竞争对手进行一些研究。我想要调查的第一件事是是否已经类似的竞争性书籍,因此我可以要求 ChatGPT 生成一份内容相同的现有书籍列表:

图 5.24 – ChatGPT 提供竞争性书籍列表的示例
你还可以就你想要出版的市场饱和度征求反馈:

图 5.25 – ChatGPT 关于如何在市场中保持竞争力的建议
最后,让 ChatGPT 更精确地告诉我为了在我运营的市场中保持竞争力我应该做什么:

图 5.26 – ChatGPT 如何建议改进书籍内容以使其脱颖出
ChatGPT 在列出一些让我的书籍独独特的技巧方面表现得非常好。
总的来说,ChatGPT 可以成为信息检索和竞争情报的价值助手。然而,重要的是记住知识库的截止日期是 2021 年:这意味着每当我们需要检索实时信息,或者在对今天进行竞争市场分析时,我们可能无法依靠 ChatGPT。
尽管如此,无论知识库截止日期如何,该工具仍然提供了良好的建议和最佳实践。
## 总结
我们在本章中看到的所有示例都是通过 ChatGPT 实现某些成就的简要代表。这些小技巧可以极地帮助你处理重复性任务(例如使用类似的模板回复邮件)或繁重任务(例如搜索背景文档或竞争情报)。
在下一章中,我们将深入探讨 ChatGPT 改变游戏的三个主要领域——开发、营销和研究。
# 5 使用 ChatGPT 开发未来
## 加入我们的 Discord 社区

[`packt.link/EarlyAccess`](https://packt.link/EarlyAccess)
在本章中,我们将讨论开发者如何利用 ChatGPT。本章重点关注 ChatGPT 在开发者领域的主要用例,包括代码审查和优化、文档生成以及代码生成。本章将提供示例并让你能够亲自尝试这些提示词。
在对为什么开发者应该利用 ChatGPT 作为日常助手进行通用介绍后,我们将关注 ChatGPT 以及它如何完成以下任务:
* 为什么开发者需要 ChatGPT?
* 生成、优化和调试代码
* 生成代码文档并调试你的代码
* 解释 `机器学习` (`ML`) 模型,帮助数据科学家和业务用户理解模型解释性
* 翻译不同的编程语言
在本章结束时,你将能够利用 ChatGPT 进行编码活动,并将其作为编码生产力的助手。
## 为什么开发者需要 ChatGPT?
我个人认为,ChatGPT 最令人惊讶的能力之一就是处理代码。处理任何类型的代码。我们已经在之前的章节中看到了一些 ChatGPT 生成 Python 代码的示例。然而,ChatGPT 对开发者而言的能力远超这个例子。它可以作为代码生成、解释和调试的日常助手。
在最流行的语言中,我们当然提到 Python、JavaScript、SQL 和 C#。然而,正如 ChatGPT 自己所言,它涵盖了广泛的语言:

图 6.1 – ChatGPT 列出了它能够理解和生成的编程语言
无论你是后端/前端开发者、数据科学家还是数据工程师,只要你使用编程语言,ChatGPT 都能成为游戏规则的改变者,我们将在接下来几节的示例中看到这一点。
从下一节开始,我们将深入探讨 ChatGPT 在处理代码时可以实现的具体示例。我们将看到涵盖不同领域的端到端用用例,以便我们熟悉如何将 ChatGPT 作为代码助手。
## 生成、优化和调试代码
你应该利用的核心能力是 ChatGPT 的代码生成。你有多少次在寻找一段预构建的代码作为起点呢?例如生成 `utils` 函数、示例数据集、`SQL` 架构等?ChatGPT 能够根据自然语言输入生成代码:

图 6.2 – ChatGPT 生成写入 CSV 文件的 Python 函数的示例
如你所见,ChatGPT 不仅能够生成该函数,还能解释该函数的功能、如何使用它,以及用诸如 `my_folder` 的通用占位符替换哪些内容。
现在让我们提高难度。如果 ChatGPT 能够生成一个 Python 函数,它还能生成整个视频游戏吗?让我们试一下。我想做给 ChatGPT 提供一张我想要开发的游戏类型的插图,并要求它用代码来实现。以下是我理想游戏的插图(你能猜出它的名字吗?):

图 1:吃豆人游戏的插图。
现在让我们要求 ChatGPT 重现它:

图 2:ChatGPT 生成 HTML、CSS 和 JS 代码的示例。
正如 ChatGPT 所言,完整游戏需要大量代码,但让我们看看生成的代码目前是如何运作的(为了运行代码,我使用了在线工具 `codepen.io`):

图 3:ChatGPT 生成的吃豆人游戏。
如你所见,原型产品看起来已经与我的目标相似了!这是一个生成式 AI 如何帮助你克服“从零开始的魔咒”的例子:事实上,从一份白皮书开始有时会令人受阻,而拥有草案产品作为起点不仅可以加快整个过程,还能激发创意并提高结果质量。
ChatGPT 也是代码优化的出色助手。事实上,它可以通过根据我们的输入创建优化脚本,从而节省一些运行时间或计算能力。这种能力在自然语言领域,可以与我们在第 5 节“提高写作技巧和翻译”部分看到的写作辅助功能相比。
例如,假设你想从另一个列表创建一个奇数列表。为了实现这一结果,你写了以下 Python 脚本(为了练习,我还将使用 `timeit` 和 `datetime` 库来跟踪执行时间):
```python
from timeit import default_timer as timer from datetime import timedelta
start = timer()
elements = list(range(1_000_000)) data = []
for el in elements: if not el % 2:
# 如果奇数 data.append(el)
end = timer() print(timedelta(seconds=end-start))
让我们看看运行需要多长时间:

图 4:Python 函数的执行速度。
执行时间为 00.115022 秒。如果我们要求 ChatGPT 优化这个脚本会发生什么?

图 6.6 – ChatGPT 为 Python 脚本生成优化替代方案
ChatGPT 为我提供了两个示例,以更低的执行时间实现相同的结果。
让我们在 Jupyter Notebook 中测试两者:
2023-03-25 11:27:10.270 Uncaught app exception Traceback (most recent call):
File "C:\Users\vaalt\Anaconda3\lib\site-packages\streamlit\runtime\ scriptrunner\script_runner.py", line 565, in _run_script
exec(code, module. dict )
File "C:\Users\vaalt\OneDrive\Desktop\medium articles\llm.py", line 129, in <module>
user_input = get_text()
File "C:\Users\vaalt\OneDrive\Desktop\medium articles\llm.py", line 50, in get_text
当你处理新的应用程序或项目时,将代码与文档联系起来是一个良好的习惯。它可能以 docstring(文档字符串)的形式存在,你可以将其嵌入到函数或类中,以便他人可以在开发环境中直接调用它们。
例如,让我们考虑上一个节中开发的同一个函数,并将其转换为一个 Python 类:
class UnderscoreAdder:
def __init__(self, word):
self.word = word
def add_underscores(self):
new_word = ""
for i in range(len(self.word)):
new_word += self.word[i] + "_"
return new_word
我们可以对其进行如下测试:

图 5:测试 UnderscoreAdder 类。
现在,假设我想要通过 UnderscoreAdder 约定来检索 docstring 文档。通过对 Python 包、函数和方法进行这种操作,我们可以获得该特定对象能力的完整文档,如下(使用 pandas Python 库的示例):
图 6.13 – pandas 库文档示例。
现在,让我们让 ChatGPT 为我们的 UnderscoreAdder 类产生相同的结果。

图 6.14 – ChatGPT 用文档更新代码。
结果,如果我们按照前面的代码更新了我们的类和 UnderscoreAdder?,我们将得到如下输出:

图 6.15 – 新的 UnderscoreAdder 类文档。
最后,还可以利用 ChatGPT 用自然语言解释脚本、函数、类或其他类似内容的功能。我们已经看到许多 ChatGPT 用清晰的解释来丰富其代码相关的回答的示例。然而,我们可以通过就代码理解方面提出特定问题来增强这种能力。
例如,让我们让 ChatGPT 向我们解释以下 Python 脚本的作用:

图 6.16 – ChatGPT 解释 Python 脚本的示例。
代码可解释性也可以是上述文档的一部分,也可以用于开发者之间,他们可能想更好地理解其他团队的复杂代码,或者(有时在我身上发生)记住他们以前写过的内容。
得益于 ChatGPT 和本节提到的功能,开发者可以轻松用自然语言跟踪项目生命周期,以便让新团队成员和非技术用户都更容易理解目前完成的工作。
我们将在下一节中看到,代码可解释性是数据科学项目中 ML 模型可解释性的关键一步。
理解 ML 模型可解释性
模型可解释性是指人类理解 ML 模型预测背后逻辑的难易程度。本质上,它是理解模型如何做出决策以及哪些变量对其预测做出贡献的能力。
让我们来看使用深度学习卷积神经网络(CNN)进行图像分类的模型可解释性示例。我使用 Python 和 Keras 构建了我的模型。为此,我将直接从 keras.datasets 下载 CIFAR-10 数据集:它包含 60,000 张 32x32 的彩色图像(即 3 通道图像),分为 10 类(飞机、汽车、鸟、猫、鹿、狗、蛙、马、船和卡车),每类 6,000 张图像。在这里,我将仅分享主体;你可以在书籍 GitHub 仓库中找到数据准备和预处理的所有相关代码:github.com/PacktPublishing/The-Ultimate-Guide-to-ChatGPT-and-OpenAI/tree/main/Chapter%206%20-%20ChatGPT%20for%20Developers/code:
model=tf.keras.Sequential()
model.add(tf.keras.layers.Conv2D(32,kernel_ size=(3,3),activation='relu',input_shape= (32,32,1)))
model.add(tf.keras.layers.MaxPooling2D(pool_size=(2,2))) model.add(tf.keras.layers.Flatten()) model.add(tf.keras.layers.Dense(1024,activation='relu')) model.add(tf.keras.layers.Dense(10,activation='softmax'))
前面的代码由执行不同操作的几个层组成。我可能对模型结构的解释以及每一层的作用感兴趣。让我们让 ChatGPT 在这方面提供帮助(在下方你可以看到回复的片段):

图 6.17 – 使用 ChatGPT 进行模型可解释性。
如前面的图所示,ChatGPT 能够为我们的 CNN 结构和层提供清晰的解释。它还会添加一些注释和提示,例如使用池化层有助于减少输入的维度。
我也可以通过 ChatGPT 在验证阶段支持解释结果。因此,在将数据分为训练集和测试集并在训练集上训练模型后,我想查看在测试集上的性能:

图 6.18 – 标。
让我们也让 ChatGPT 详细说明我们的验证指标(截断输出):

图 6.19 – ChatGPT 解释指标的示例。
再次,结果非常令人印象深刻,它就如何针对训练和测试集设置 ML 实验提供了清晰的指导。
模型可解释性重要有很多原因。一个关键因素是它缩小了业务用户与模型代码之间的差距。这是让业务用户理解模型如何运行以及将其转化为代码想法的关键。
此外,模型可解释性实现了负责任和伦理 AI 的核心原则之一,即 AI 系统背后思考和行为的透明度。解锁模型可解释性意味着检测模型在运行时可能出现的偏见或有害行为,并防止它们的发生。
总而言之,正如我们在前面的示例所示,ChatGPT 在模型可解释性方面可以提供价值支持。
我们将探索的下一个 ChatGPT 功能是开发者生产力的另一个助力,当同一个项目使用多种编程语言时。
不同编程语言之间的翻译
在第 5 节中,我们看到了 ChatGPT 在不同语言之间翻译的强大能力。真正令人惊的是,自然语言并不是它唯一的翻译对象。事实上,ChatGPT 能够在不同的编程语言之间进行翻译,同时保持相同的输出和风格(即保留存在的 docstring 文档)。
在许多场景下,这可能会改变游戏规则。
例如,你可能需要学习一种以前从未见过的新编程语言或统计工具,因为你需要快速交付该项目。在 ChatGPT 的帮助下,你可以用你偏好的语言开始编程,然后让它翻译成目标语言,你可以在翻译的过程中同时学习该语言。
想象一下,某个项目需要在 MATLAB(一种由 MathWorks 开发的专有数值计算和编程软件)中交付,而你一直都在用 Python 进行编程。该项目包含对美国国家标准与技术研究院(MNIST)数据集中的图像进行分类(原始数据集描述和相关论文可以在此处通过 yann.lecun.com/exdb/mnist/ 找到)。该数据集包含大量手写数字,经常被用于教授各种图像处理系统。
首先,我编写了以下 Python 代码来初始化一个用于分类的深度学习模型:
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers
## Load the MNIST dataset
(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_ data()
# Preprocess the data
x_train = x_train.reshape(-1, 28*28) / 255.0 x_test = x_test.reshape(-1, 28*28) / 255.0 y_train = keras.utils.to_categorical(y_train) y_test = keras.utils.to_categorical(y_test)
# Define the model architecture model = keras.Sequential([
layers.Dense(256, activation='relu', input_shape=(28*28,)), layers.Dense(128, activation='relu'),
layers.Dense(10, activation='softmax')
])
# Compile the model
model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
# Train the model
history = model.fit(x_train, y_train, validation_split=0.2, epochs=10, batch_size=128)
# Evaluate the model
test_loss, test_acc = model.evaluate(x_test, y_test, verbose=0) print('Test accuracy:', test_acc)
现在让我们看看如果我们将前面的代码作为上下文提供给 ChatGPT 并让它将其翻译成 MATLAB 会发生什么:

图 20 – ChatGPT 将 Python 代码翻译为 MATLAB
让我们看看它是否也能将其翻译成其他语言,例如 JavaScript(注意,我不需要重复代码,因为代码已经在对话上下文中了):

图 6.21 – ChatGPT 将 Python 代码翻译为 JavaScript
代码翻译还可以缩小新技术与当前编程能力之间的技能差距。
代码翻译的另一个关键意义是应用现代化。事实上,想象你想要刷新你的应用栈,即迁移到云端。你可以决定从一个简单的“平移并迁移”(lift and shift)开始,迈向即服务服务(IaaS)实例(例如 Windows 或 Linux 虚拟机(VMs))。然而,在第二阶段,你可能希望对你的应用程序进行重构、重新架构或彻底重建。
下面的图表描述了应用现代化的各种选项:

图 6.22 – 你可以将应用程序迁移到公云的四种方法
ChatGPT 和 OpenAI Codex 模型可以帮助你完成迁移。以大型机为例。
大型机是大型组织主要用于执行关键任务的计算机,例如用于人口普查、消费者和行业统计、企业资源计划以及大规模交易处理的批量数据处理。大型机环境的应用编程语言是通用面向业务语言(COBOL)。尽管 COBOL 发明于 1959 年,但它在今天仍在在使用,是世界上最古老的编程语言之一。
随着技术的不断进步,位于大型机领域的应用程序一直处于持续的迁移和现代化过程中,旨在从接口、代码、成本、性能和可维护性等领域增强现有的大型机基础设施。
当然,这意味着要把 COBOL 翻译成更现代的编程语言,如 C# 或 Java。问题是,大多数新一代程序员并不熟悉 COBOL;因此,在这种背景下存在巨大的技能差距。
让我们考虑一个 COBOL 脚本,它读取一个输入数字,对其加 10,然后打印结果。
IDENTIFICATION DIVISION.
PROGRAM-ID. AddTen.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 INPUT-NUMBER PIC 9(5).
01 RESULT-NUMBER PIC 9(5).
PROCEDURE DIVISION.
DISPLAY 'Enter a number: '.
ACCEPT INPUT-NUMBER.
COMPUTE RESULT-NUMBER = INPUT-NUMBER + 10.
DISPLAY 'Result after adding 10: ' RESULT-NUMBER.
STOP RUN.
然后我将之前的 COBOL 脚本传递给 ChatGPT,以便它可以将其作为上下文来制定其回复。现在让我们让 ChatGPT 将该脚本翻译成 C#:

图 6.23 – ChatGPT 将 COBOL 翻译为 C# 的示例
诸如 ChatGPT 的工具可以通过引入一个既了解编程过去又未来的层,帮助缩小此类及类似场景中的技能差距。
结论,ChatGPT 可以作为应用现代化的有效工具,在提供代码升级的同时,提供宝贵的见解和建议。
这是一个示例。ChatGPT 在营销领域最突出且最有前景的用例之一是个性化营销。ChatGPT 可用于分析客户数据,并生成能与单个客户产生共鸣的个性化营销信息。例如,营销团队可以使用 ChatGPT 分析客户数据,并开发针对特定客户偏好和行为的定制电子邮件营销活动。这可以提高转化概率,并带来更高的客户满意度。通过提供对客户情感和行为的洞察、生成个性化营销信息、提供个性化客户支持以及生成内容,ChatGPT 可以帮助营销人员提供卓越的客户体验并推动业务增长。
这是 ChatGPT 在营销中许多应用示例中的之一。在接下来的章节中,我们将查看由 ChatGPT 支持的端到端营销项目的具体示例。
新产品开发与市场进入策略
你将 ChatGPT 引入营销活动的第一种方式可能是新产品开发和 go-to-market (GTM)策略的助手。
在本节中,我们将介绍一个如何开发和推广新产品的逐步指南。你已经拥有一个名为 RunFast 的跑步服装品牌,到目前为止只生产了鞋子,因此你想通过新的产品线来扩大业务。我们将从头脑风暴创意来创建 GTM 策略开始。当然,一切都得到了 ChatGPT 的支持:
- 头脑风暴创意:ChatGPT 能支持你的的第一件事是为你新产品线进行头脑风暴并起草方案。它还会提供每个建议背后的逻辑。所以,让我们问问我应该关注什么样的新产品线:

图 7.1 – ChatGPT 生成的新创意示例
在三个建议中,我们将选择第二个,因为我们可以借此对环境产生积极影响,并提升我们的品牌声誉。具体来说,我将从环保跑步袜开始。
- 产品名称:现在我们已经确定了主意,我们需要为它想一个响亮的名字。同样,我会要求 ChatGPT 提供更多选项,以便我从中选出最喜欢的:

图 7.2 – 潜在产品名称列表
GreenStride 听起来不错——我就用这一个。
- 生成响亮的口号:除了产品名称,我还想分享名称背后的意图和产品线的使命,以便吸引我的目标受众。我想激发客户的信任和忠诚度,让他们在我的新产品线背后的使命中看到自己。

图 7.3 – 我的新产品名称的口号列表
太棒了——现在我对稍后将使用的产品名称和口号感到满意,并将用它们创建一个独特的社交媒体公告。在在此之前,我想花更多时间对目标受众进行市场研究。

图 7.4 – 我的新产品线要触达的目标群体列表
考虑到受众中不同的群体是重要的,这样你就可以区分你想要传达的信息。在我的案例中,我想确保我的产品线能够涵盖不同的人群,例如竞技性跑步者、休闲跑步者和健身爱好者。
- 产品变体与销售渠道:根据上述的潜在客户群体,我可以生成产品变体,使其更针对特定受众:

图 7.5 – 产品线变体示例
同样,我也可以要求 ChatGPT 为上述每个小组建议不同的销售渠道:

图 7.6 – ChatGPT 对不同销售渠道的建议
- 竞争脱颖而出:我希望我的产品线能在竞争中脱颖而出,并在饱和的市场中出现——我想让它具有独特性。带着这个目标,我要求 ChatGPT 包含社会因素,如可持续性和包容性。让我们就这方面向 ChatGPT 要建议:

图 7.7 – ChatGPT 生成的卓越特性示例
如你所见,它能够生成让我的产品线变得有趣的特性。
- 产品描述:现在是时候开始构建市场进入(Go-To-Market)计划了。首先,我想为我的网站生成一个产品描述,包括之前所有的独特差异化点。

图 7.8 – ChatGPT 生成的描述和 SEO 关键词示例
- 公平价格:另一个关键要素是为我们的产品确定公平的价格。由于我针对不同受众(竞技性跑步者、休闲跑步者和健身爱好者)区分了产品变体,我也希望有一个考虑到这种分群的价格范围。注意,在接下来的示例中,ChatGPT 正在调用网络搜索插件来检索有关当前跑步袜市场定价更新信息。

图 7.9 – 产品变体的价格范围
我们快完成了。我们已经经历了许多新产品开发和市场进入的步骤,在每一步中,ChatGPT 都充当了强大的支持工具。
作为最后一件事,我们可以要求 ChatGPT 生成一个关于我们新产品的 Instagram 帖子,包括相关的标签和 SEO 关键词。然后通过 DALL-E 生成图像,它是 ChatGPT Plus 中的内置插件。

图 7.10 – ChatGPT 生成的社交媒体帖子
以及,通过 DALL-E 的特别贡献:

图 1:由 DALL-E 3 支持 生成的插图示例
这就是最终的结果:

图 7.11 – 由 ChatGPT 和 DALL-E 完全生成的 Instagram 帖子
当然,对于完整的产品开发和市场进入,这里缺少许多元素。然而,通过 ChatGPT 的支持(以及 DALL-E 3 的特别贡献),我们成功地头脑风暴了新的产品线和变体、潜在客户、响亮的口号,并生成了一个非常漂亮的 Instagram 帖子来宣布 GreenStride 的发布!
用于营销对比的 A/B 测试
营销中的 A/B 测试是一种比较不同版本的营销活动、广告或网站的方法,以确定哪一个表现更好。在 A/B 测试中,会创建同一个营销活动或元素的两个版本,两个版本之间只改变一个变量。目标是看哪个版本能产生更多的点击、转化或其他理想结果。
A/B 测试的一个示例可能是测试两个版本的邮件营销活动,使用不同的主题行;或者测试两个版本的网站着陆页,使用不同的呼吁行动(call-to-action)按钮。通过衡量每个版本的响应率,营销人员可以确定哪个版本表现更好,并就未来使用哪个版本做出数据驱动的决策。
A/B 测试允许营销人员优化其活动和元素以实现最大效能,从而获得更好的结果和更高的投资回报率(ROI)。
由于这种方法涉及生成同一份内容的许多变体的过程,ChatGPT 的生成能力绝对可以为此提供帮助。
让我们考虑以下例子。我正在推广我开发的一个新产品:一种为快速攀岩者设计的全新、轻薄型攀岩安全带。我已经做了一些市场研究,了解我的小众受众。我 known 与该受众交流的一个绝佳渠道是在在线攀岩博客上发布内容,大多数攀岩馆的成员都是该博客的读者。
我的目标是创建一个优秀的博客文章来分享这款新款安全带的发布,我想在两个组中测试它的两个不同版本。我正准备发布并希望作为 A/B 测试对象的博客文章如下:

图 7.12 – 发布攀岩装备的博客文章示例
在这里,ChatGPT 可以在两个层面上帮助我们:
- 第一个层面是重新措述文章,使用不同的关键词或不同的吸引人球的口号。为了做到这一点,一旦此帖子作为上下文提供,我们可以要求
ChatGPT对文章进行处理并轻微更改某些元素:

图 7.13 – ChatGPT 生成的博客文章新版本
根据我的请求,ChatGPT 能够仅重新我请求的那些元素(标题、副标题和结束句子),以便我通过监控两个受众组的反应来监控这些元素的效效。
- 第二个层面是进行网页设计,即更改图像的搭配而不是按钮的位置。为了这个目的,我为发布在攀岩博客上的博文创建了一个简单的网页(你可以在书的 GitHub 仓库中找到代码,地址为 [
https://github.com/PacktPublishing/The-Ultimate-Guide-to-ChatGPT-and-OpenAI/tree/main/Chapter%207%20-%20ChatGPT%20for%20Marketers/Code](https://github.com/PacktPublishing/The-Ultimate-Guide-to-ChatGPT-and-OpenAI/tree/main/Chapter 7 - ChatGPT for Marketers/Code]`):

图 7.14 – 发布在攀岩博客上的示例博文
我们可以直接向 ChatGPT 馈送 HTML 代码并要求它更改某些布局元素,例如按钮的位置或它们的措辞。例如,比起“立即购买”(Buy Now),读者可能被“我想要!”(I want one!)按钮更吸引。
所以,让我们给 ChatGPT 提供 HTML 源代码:

图 7.14 – ChatGPT 修改 HTML 代码
让我们看看输出是什么样的(我也将标题、副标题和段落更改为 ChatGPT 生成的内容):

图 7.15 – 网站的新版本
如你所示,ChatGPT 只在按钮层面进行了干预,轻微改变了它们的布局、位置、颜色和措辞。
结论,ChatGPT 是营销 A/B 测试中非常有价值的工具。它快速生成相同内容的不同版本的能力可以缩短新活动的上市时间。通过利用 ChatGPT 进行 A/B 测试,你可以优化营销策略,并最终为业务带来更好的结果。
提升搜索引擎优化 (SEO)
ChatGPT 另一个可能改变规则的光前景领域是搜索引擎优化(SEO)。这是在 Google 或 Bing 等搜索引擎中获得排名的关键元素,它决定了你的网站是否对那些寻找你所推广内容的用户可见。
SEO 是一种用于提高网站在搜索引擎结果页面页(
SERPs)上可见性和排名的技术。它是通过优化网站或网页来增加来自搜索引擎的自然(非付费)流量的数量和质量来实现的。SEO 的目的是通过针对特定关键词或短语优化网站来吸引更多有针对性的访问者。
想象你经营着名为 Hat&Gloves 的电子商务公司,正如你可能猜的那样,它只卖帽子和手套。你现在正在创建你的电子商务网站并想优化其排名。让我们 ChatGPT 列出一些嵌入我们网站的 {关键词:

图 7.17 – ChatGPT 生成的 SEO 关键词示例
如你所示,ChatGPT 能够创建一个不同类型的关键词列表。其中一些非常直观,例如“帽子和手套”(Hats and Gloves)。其他一些是相关的,有间接联系。例如,“礼物建议”(Gift ideas)并不一定与我的电子商务业务相关,但将其包含进去可能是明智的,这样我可以扩大我的受众。
最后,我们还可以进一步利用“扮演……”(Act as…)这种技巧,这我们在第 3 章中提到过。对我们的网站进行评估并理解其是否已按预期进行了优化,将是非常有趣的事情。在营销领域,这种分析被称为 SEO audit(SEO 审计)。SEO 审计是对网站 SEO 性能和潜在改进领域的评估。SEO 审计通常由 SEO 专家、网页开发人员或营销人员进行,涉及对网站的技术架构、内容和反链配置的全面分析。
在SEO 审计过程中,审计人员通常会使用一系列工具和技术来识别改进领域,例如关键词分析、网站速度分析、网站架构分析和内容分析。随后,审计人员将生成报告,列出关键问题、改进机会以及解决这些问题的建议行动。
让我们让 ChatGPT 扮演 SEO 专家来执行此审计。我们将使用上面提到的攀岩博客作为参考网站。我将给 ChatGPT 代码并要求其执行以下指令:“扮演 SEO 专家并对上述 HTML 代码生成一份简短的 SEO 审计(最多 300 字)”。这是回复:

图 7.19 – ChatGPT 对攀岩博客 HTML 代码生成 SEO 审计。
ChatGPT 能够生成相当准确的分析,并带有相关的评论和建议。总的来说,ChatGPT 在 SEO 相关活动方面具有有趣的潜力,无论你是从零开始构建网站还是想改进现有网站,它都是一个很好的工具。
通过情感分析提高质量并增加客户满意度
情感分析是营销中一种技术,用于分析和解释客户对品牌、产品或服务所表达的情绪和观点。它涉及使用natural language processing(NLP)和machine learning(ML) 算法来识别和分类社交媒体帖子、客户评论和反馈调查等文本数据的情感。
通过执行情感分析,营销人员可以洞察客户对其品牌的感知,识别改进领域,并做出数据驱动的决策以优化营销策略。例如,他们可以跟踪客户评论的情绪,识别哪些产品或服务收到了正面或负面反馈,并相应调整他们的营销信息。
总而言,情感分析是营销人员理解客户情感、衡量客户满意度以及开发与目标受众产生共鸣的有效营销活动的宝贵工具。
情感分析已经存在一段时间了,所以你可能好奇 ChatGPT 能带来什么额外价值。嗯,除了分析的准确性(它是目前市场上最强大的模型)之外,ChatGPT 与其他情感分析工具的区别在于它是general artificial intelligence(AGI,通用人工智能)。
这意味着当我们使用 ChatGPT 进行情感分析时,我们并没有使用它用于该任务的特定 API:ChatGPT 和 OpenAI 模型背后的核心思想是它们可以同时协助用户处理许多通用任务,与任务交互并根据用户的请求更改分析的范围。
因此,ChatGPT 确实能够捕捉给定文本(如推特帖子或产品评论)的情绪。然而,ChatGPT 还可以进一步协助识别产品或品牌中对情感产生正面或负面影响的特定方面方面。例如,如果客户不断以负面方式提到产品的某个功能,ChatGPT 可以将该功能突出为待改进领域。或者,ChatGPT 可能被要求对一条特别微妙的评论生成回复,同时兼顾评论的情绪并将其作为回复的背景。同样,它可以生成报告,总结评论或留言中发现的所有负面和正面元素,并将它们归入不同的类别。
让我们考虑以下示例。一位客户最近从我的电子商务公司 RunFast 购买了一双鞋,并留下了以下评论:
“我最近买了 RunFast Prodigy 鞋子,感觉很复杂。它们非常舒适,缓冲和支撑非常好,减轻了我跑步时的脚部疲劳。设计也很吸引人,我收到了一些赞美。然而,耐用性令人失望;外底磨损很快,透气鞋面在几周后就有磨损迹象。考虑到高价格,尽管它们舒适且漂亮,我还是不推荐。”
让我们让 ChatGPT 捕捉这条评论的情绪:

图 7.20 – ChatGPT 分析客户评论
从前面的图中,我们可以看到 ChatGPT 并不限于提供一个标签:它还解释了构成该评论的正面和负面元素,这条评论具有复杂感,因此整体上可以被标记为中性。
让我们尝试深入探讨,并就改进产品提出一些建议:

图 7.21 – 根据客户反馈如何改进我产品的建议
最后,让我们给客户生成回复,表明我们作为一家公司关心客户的反馈,并希望改进我们的产品。

图 7.22 – ChatGPT 生成的回复
我们看到的示例非常简单,只有一条评论。现在想象一下我们有成千上万条评论,以及接收反馈的多种销售渠道。想象一下 ChatGPT 和 OpenAI 模型等工具的力量,它们能够分析并整合所有这些信息,识别产品的优缺点,捕捉客户趋势和购物习惯。此外,对于客户服务和留存,我们还可以使用我们偏好的写作风格自动回复评论。事实上,通过调整机器人的语言和语调以满足客户的特定需求和期望,你可以创造更具吸引力且有效的客户体验。
以下是一些示例:
-
共情机器人:使用同理心的语气和语言与可能遇到问题或需要敏感问题帮助的客户进行交互。 -
专业机器人:使用专业的语气和语言与可能寻找特定信息或需要技术问题帮助的客户进行交互。 -
对话机器人:使用随和和友好的语气与可能寻求个性化体验或有更一般查询的客户进行交互。 -
幽默机器人:使用幽默和机智的语言与可能寻求轻松体验或缓解紧张局势的客户进行交互。 -
教育机器人:使用教学式的沟通风格与可能想要了解更多产品或服务的客户进行交互。
总之,ChatGPT 可以成为企业进行情感分析、提高质量和留住客户的强大工具。凭借先进的自然语言处理能力,ChatGPT 可以实时准确地分析客户反馈和评论,为企业提供关于客户情绪和偏好的宝贵见解。通过将 ChatGPT 作为客户体验策略的一部分,企业可以快速识别任何可能对客户满意度产生负面影响的问题并采取纠正措施。这不仅可以帮助企业提高质量,还能增加客户的忠诚度和留存率。
总结
在本章中,我们探索了营销人员如何利用 ChatGPT 来增强其营销策略。我们了解到 ChatGPT 可以帮助开发新产品,以及定义其进入市场的策略、设计 A/B 测试、增强 SEO 分析,并捕捉评论、社交媒体帖子和其他客户反馈的情绪。
ChatGPT 对营销人员的重要性在于,它具有改变公司与客户互动方式的潜力。通过利用 NLP、ML 和大数据的力量,ChatGPT 允许公司创建更具个性化和相关性的营销信息,提高客户支持和满意度,最终驱动销售和收入。
随着 ChatGPT 的不断进步和演进,我们可能会看到它更多地参与营销行业,特别是在公司与客户互动的方式上。事实上,高度依赖 AI 允许公司对客户行为和偏好获得更深层次的见解。
营销人员的关键启示是拥抱这些变化并适应 AI 驱动的营销的新现实,以保持竞争领先地位并满足客户的需求。
在下一章中,我们将介绍书中涵盖的 ChatGPT 应用的第三个也是最后一个领域——研究。
第 7 章 用 ChatGPT 重塑研究
加入我们的 Discord 书籍社区

在本章中,我们重点关注希望利用 ChatGPT 的研究人员。本章将阐述 ChatGPT 可以解决的几个主要用例,以便你通过具体的示例学习如何在研究中使用 ChatGPT。
在本章结束时,你将熟悉熟悉如何以多种方式将 ChatGPT 作为研究助手,包括以下:
-
研究人员对 ChatGPT 的需求
-
为你的研究头脑风暴文献
-
为你的实验设计和框架提供支持
-
生成并格式化参考文献,以将其整合到你的研究中
-
针对不同受众交付关于你的研究推演或幻灯片演示
本章还将提供示例,并让你能够亲自尝试提示词(prompts)。
研究人员对 ChatGPT 的需求
ChatGPT 是各个领域研究人员极具价值的资源。作为一个在海量数据上训练的复杂语言模型,ChatGPT 可以快速准确地处理大量信息,并生成可能通过传统研究方法难以发现或耗时才能发现的见解。
此外,ChatGPT 可以通过分析人类研究人员可能无法立即察觉的模式和趋势,为研究人员提供其领域的独特视角。例如,想象一位研究人员正在研究气候变化,并希望理解公众对这一问题的看法。他们可能会要求 ChatGPT 分析与气候相关的社交媒体数据,并识别网上人们表达的最常见主题和情绪。然后,ChatGPT 可以为研究人员提供一份详尽报告,列出与该主题相关的常见词汇、短语和情感,以及任何可能需要了解的新趋势或模式。
通过与 ChatGPT 合作,研究人员可以获得端技术和见解,并保持在该领域的前沿。
现在让我们深入探讨 ChatGPT 提高研究效率的四个用例。
本章中提出的大多数示例都是基于最新信息的;事实上,你经常会看到 ChatGPT 调用网络搜索插件。如果你使用的是带有
GPT-3.5-turbo的 ChatGPT(即免费版),请记住它没有启用网络搜索插件,因此其知识截止日期限制在 2021 年。如果你正在寻找更新的信息,以及通常来自网络的引用(以便你可以双击回答的可靠性),这可能是一个局限性。
为你的研究头脑风暴文献
文献综述是一个关键且系统的过程,旨在对特定主题或问题的现有已发表研究进行检查。它涉及搜索、评阅和综合相关的已发表研究和其他来源(如书籍、会议论文和灰色文献)。文献综述的目标是识别特定领域中的空白、不一致之处以及开展进一步研究的机会。
文献综述过程通常涉及以下步骤:
- 定义研究问题:进行文献综述的第一步是定义感兴趣主题的研究问题。假设我们正在进行关于社交媒体对心理健康影响的研究。现在我们想对头脑风暴一些可能的研究问题以集中我们的研究,我们可以利用 ChatGPT 来做到这一点:

图 8.1 – 基于给定主题的研究问题示例
这些都是可以进一步调查的有趣问题。由于我特别对第一个感兴趣——“社交媒体互动和在线支持社区以哪些方式影响患有慢性心理健康状况个人的心理健康?”——我将保留这作为我们分析后续步骤的参考。
- 搜索文献:现在我们有了研究问题,下一步是使用各种数据库、搜索引擎和其他来源搜索相关文献。研究人员可以使用特定的关键词和搜索术语帮助识别相关研究。

图 8.2 – 在 ChatGPT 支持下的文献检索
从 ChatGPT 的建议开始,我们可以深入研究这些参考文献。
- 筛选文献:一旦确定了相关文献,下一步是对研究进行筛选,以确定它们是否符合综述的标准。这通常涉及审阅摘要,并在必要时审阅研究全文。
假设我们想深入研究《社交媒体与青少年焦虑:来自哈佛大学教育学院的见解》研究论文。让我们让 ChatGPT 为我们筛选:

图 8.3 – 特定论文的筛选
ChatGPT 能够为我提供该论文的概述,考虑到其研究问题和讨论的主要主题,我认为它对我自己的研究非常有用。
- 提取数据:在确定了相关研究后,研究人员需要从每项研究中提取数据,例如研究设计、样本量、数据收集方法和主要发现。
例如,假设我们从 Hinduja 和 Patchin (2018) 年的《数字自残:流行率、动机和结果》论文中收集以下信息:
-
论文中收集的数据来源和研究对象
-
研究者采用的数据收集方法
-
数据样本量
-
分析的主要局限性和缺点
-
研究者采用的实验设计
具体如下:

图 8.4 – 从给定论文中提取相关数据和框架
1. 文献综合
文献综述过程的最后一步是综合研究结果,并对该领域的当前知识状态得出结论。这可能涉及识别共同主题、突出文献中的空白或不一致之处,以及确定未来研究的机会。
让我们想象一下,除了 ChatGPT 提出的论文外,我们还收集了其他想要综合的标题和论文。更具体地说,我想了解它们是否得出了相同的结论、共同趋势是什么,以及哪种方法比其他方法更可靠。对于这种场景,我们将考虑三篇研究论文:
-
The Effects of Social Media on Mental Health: A Proposed Study,作者为 Grant Sean Bossard (
digitalcommons.bard.edu/cgi/viewcontent. cgi?article=1028&context=senproj_f2020) -
The Impact of Social Media on Mental Health,作者为 Vardanush Palyan (
www. spotlightonresearch.com/mental-health-research/the-impact-of- social-media-on-mental-health) -
The Impact of Social Media on Mental Health: a mixed methods research of service providers’ awareness,作者为 Sarah Nichole Koehler 和 Bobbie Rose Parrell (
scholarworks. lib.csusb.edu/cgi/viewcontent.cgi?article=2131&context=etd)
结果显示如下:

图 8.5 – 三篇研究论文的文献分析与基准测试
此外,在这种情况下,ChatGPT 能够对提供的三篇论文生成相关的总结和分析,包括方法之间的基准测试和可靠性考虑。
总的来说,ChatGPT 能够在文献综述领域开展许多活动,从研究问题的脑脑风暴到文献综合。一如既常,循环中需要一个领域专家(SME)来审核结果,但在这种帮助下,许多活动可以更高效地完成。
另一个可以由 ChatGPT 支持的活动是研究者希望进行的实验设计。我们将在下一节介绍这一点。
为你的实验设计和框架提供支持
实验设计是计划和执行科学实验或研究以回答研究问题的过程。它涉及对研究设计、待测变量、样本量以及数据收集和分析程序的决策。
ChatGPT 可以通过向你建议研究框架(例如随机对照试验、准实验设计或相关性研究)来帮助研究的实验设计,并在设计的实施的同时提供支持。
让我们考虑以下场景。我们想调查新的教育计划对学生数学学习成果的影响。这种新计划涉及项目制学习(PBL),这意味着要求学生合作完成真实世界的项目,使用数学概念和技能来解决问题并创建解决方案。
为了实现这一目的,我们将研究问题定义如下:
新的 PBL 程序在提高学生表现方面与传统教学方法相比如何?
以下是 ChatGPT 如何提供帮助:
- 确定研究设计:ChatGPT 可以协助为研究问题确定合适的研究设计,例如随机对照试验、准实验设计或相关性研究。

图 8.6 – ChatGPT 为你的实验建议合适的实验设计
ChatGPT 建议采用随机对照试验 (RCT),并对其背后的原因提供了清晰的解释。对我来说,采用这种方法是合理的:下一步将确定实验中考虑的结果衡量和变量。
- 识别结果衡量:ChatGPT 可以帮助你识别一些潜在的结果衡量,以确定测试结果。让我们为我们的研究寻求一些建议:

图 8.7 – 给定研究学习的学习成果
我选择测试分作为结果衡量是合理的。
- 识别变量:ChatGPT 可以帮助研究者识别研究中的自变量和因变量:

图 8.8 – ChatGPT 为给定研究生成变量
注意,ChatGPT 还能生成特定于我们考虑的研究设计(RCT)的变量类型,称为控制变量。
控制变量(也称为协变量)是在研究中保持保持不变或受控制的变量,以隔离自变量与因变量之间的关系。这些变量不是研究的重点,但将它们包含进来是为了减少混杂变量对结果的影响。通过控制这些变量,研究人员可以降低获得假阳性或假阴性结果的风险,并增加研究的内部效度。
有了上述变量,我们准备好建立实验了。现在我们需要选择参与者,ChatGPT 可以协助我们完成这一点。
- 采样策略:ChatGPT 可以为研究建议潜在的采样策略:

图 8.9 – ChatGPT 的 RCT 采样建议
注意,在现实场景中,要求 AI 工具生成带有解释的更多选项是一个良好的做法,这样你就可以做出合理的决策。对于这个例子,让我们按照 ChatGPT 给我们的建议进行,这还包括对目标群体和样本量的建议。
- 数据分析:ChatGPT 可以协助研究者确定用于分析研究收集数据的合适统计方法,例如方差分析 (ANOVA)、t 检验或回归分析。

图 8.10 – ChatGPT 为给定研究建议统计检验
ChatGPT 建议的所有内容都是连贯的,并在关于如何进行统计检验的论文中找到了证实。它还能识别出我们可能谈论的是连续变量(即分数),以便我们知道后续的所有信息都是基于这一假设的。如果我们想要离散分数,我们可以通过添加此信息来调整提示词,ChatGPT 然后就会建议不同的方法。
ChatGPT 指明假设并解释其推理,这是根据其输入做出安全决策的关键。
结论,ChatGPT 在研究人员设计实验时是一项非常有价值的工具。通过利用其自然语言处理(NLP)能力和丰富的知识库,ChatGPT 可以帮助研究人员选择合适的研究设计、确定采样技术、识别变量和学习成果,甚至建议用于分析数据的统计方法。
在下一节中,我们将进一步探索 ChatGPT 如何支持研究人员,重点关注参考文献的生成。
生成和格式化参考文献
ChatGPT 可以通过提供自动化的引用和参考文献工具来支持研究人员生成参考文献。这些工具可以为各种来源(包括书籍、文章、网站等)生成准确的引用和参考文献。ChatGPT 了解多种引用格式,例如 APA、MLA、Chicago 和 Harvard,允许研究人员根据其工作选择合适的格式。此外,ChatGPT 还可以根据研究人员的输入建议相关的来源,帮助简化研究流程并确保所有必要的来源都包含在参考文献中。通过利用这些工具,研究人员可以节省时间,并确保参考文献的准确和全面。
让我们考虑以下例子。假设我们定了一篇题为《技术对工作生产力的影响:实证研究》(The Impact of Technology on Workplace Productivity: An Empirical Study)的论文。在研究和写作过程中,我们收集了以下论文、网站、视频和其他来源的引用,我们需要将其纳入参考文献(顺序为:三篇研究论文、一个 YouTube 视频和一个网站):
-
The second machine age: Work, progress, and prosperity in a time a brilliant technologies. Brynjolfsson, 2014. https://psycnet.apa.org/record/2014-07087-000 -
The Impact of Technostress on Role Stress and Productivity. Tarafdar, 2014. Pages 301-328. https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109 -
The big debate about the future of work, explained. Vox. https://www.youtube.com/ watch?v=TUmyygCMMGA
显然,我们不能在研究论文中直接使用上述列表;我们需要对其进行正确的格式化。为了做到,我们可以向 ChatGPT 提供原始参考文献列表,并要求它以特定的格式重新生成,例如 APA 格式,这是美国心理学会(APA)的官方格式,常用于教育、心理学和社会科学的参考文献格式样式。
让我们看看 ChatGPT 是如何处理的:

图 8.11 – ChatGPT 生成的 APA 格式参考文献列表
注意,我特别指定了不要添加细节,以防 ChatGPT 不知道这些信息。事实上,我注意到 ChatGPT 有时会添加出版的月份和日期,从而导致一些错误。
ChatGPT 可以提供的其他有趣帮助是建议我们可能想要引用的潜在参考文献。我们在本章的第一段已经看到了 ChatGPT 如何在写作过程之前对相关文献进行头脑风暴;然而,一旦论文完成,我们可能会忘记引用相关文献,甚至没有意识到引用了别人的工作。
ChatGPT 在头脑风暴我们可能遗漏的潜在参考文献方面可以是一个非常出色的助手。让我们再次考虑我们的论文,研究问题是“社交媒体互动和在线支持社区如何影响患有慢性心理健康状况个人的心理健康?”假设我们设定了以下标题和摘要:其摘要包含以下内容:
Title: 社交媒体和在线支持社区对慢性心理健康状况个人心理健康的影响
Abstract: 本研究探索了社交媒体互动和在线支持社区对慢性心理健康状况个人心理健康的影响。通过检查各种在线平台,该研究旨在识别积极和消极的影响。研究利用了定性和定量方法(包括问调查和访谈)来评估焦虑、抑郁和整体生活满意度的变化。初步结果表明,虽然在线支持可以提供宝贵的社交连接和情感支持,但过度使用和负面互动可能会加剧心理健康问题。
让我们让 ChatGPT 列出可能与此类研究相关的各种参考文献:

图 8.12 – 与提供的摘要相关的参考文献列表
你也可以对论文的其他部分重复此过程,以确保你没有遗漏任何需要包含在参考文献中的相关引用。
一旦你的研究准备就绪,你可能需要通过“电梯演讲”(elevator pitch)来展示它。在下一节中,我们将看到 ChatGPT 如何也支持这项任务。
生成研究演示报告
研究的“最后一公里”通常是向不同的观众进行展示。这可能涉及准备幻灯片组、演说稿(pitch)或网络研讨会,研究人员需要面对不同类型的听众。
例如,假设我们的研究《社交媒体和在线支持社区对慢性心理健康状况个人心理健康的影响》是为了硕士论文答辩准备。在这种情况下,我们可以要求 ChatGPT 生成一个持续 15 分钟并遵循科学方法的演说结构。让我们看看产生了什么样的结果(作为背景,我指的是上段的摘要):

图 8.13 – ChatGPT 生成的论文答辩
这太令人印象深刻了!在我的大学时代,如果有这样一个工具来协助我设计答辩将非常有帮助。
从这个结构开始,我们还可以要求 ChatGPT 生成一个幻灯片组作为我们论文答辩的视觉辅助。
让我们继续这个请求:

图 8.14 – 基于演说稿的幻灯片结构
最后,假设我们的论文答辩非常出色,以至于可能获得研究资金以继续调查该课题。现在我们需要一个电梯演讲来说服委员会。让我们向 ChatGPT 请求支持:

图 8.15 – 为给定论文的电梯演讲
我们可以随时调整结果,使其更符合我们的需求,然而,现成的结构和框架可以节省大量时间,让我们能够专注于我们想要呈现的技术内容。
总的来说,ChatGPT 能够支持研究的全全程,从文献收集和综述到研究的最终演示,我们已经证明了它可以成为研究人员极出色的 AI 助手。
此外,请注意在研究领域,最近开发了一些与 ChatGPT 不同但同样由 GPT 模型驱动的工具。例如 Humanata.ai,这是一个 AI 驱动的工具,允许你上传文档并对其进行多种操作,包括摘要、即时问答以及基于上传的文件生成新论文。
这表明基于 GPT 的工具(包括 ChatGPT)正在为研究领域的若干创新铺平道路。
总结
在本章中,我们探索了 ChatGPT 作为研究人员价值工具的使用。通过文献综述、实验设计、参考文献生成和格式化以及演示报告生成,ChatGPT 可以帮助研究人员加速那些低价值或零价值的活动,让他们能够专注于相关的活动。
请注意,我们仅关注了 ChatGPT 可以支持研究人员的一小部分活动。在研究领域还有许多其他活动可以从 ChatGPT 的支持中受益,其中我们可以提到的数据收集、研究参与者招募、研究网络建立、公众参与等。
将此工具整合到工作中的研究人员可以从其多功能和节省时间的特性中受益,最终带来更有影响力的研究成果。
然而,重要的是记住,ChatGPT 仅仅是一个工具,应该结合专家知识和判断来使用。与任何研究项目一样,必须对研究问题和研究设计进行仔细考虑,以确保结果的有效性和可靠性。
通过本章,我们也结束了本书的第二部分,该部分重点关注了你可以利用 ChatGPT 的广泛场景和领域。然而,我们主要关注个人或小团队的使用,从个人生产力到研究辅助等。从第三部分开始,我们将讨论提升到大型组织如何利用 Microsoft Azure 云上提供的 OpenAI 模型 API,将 ChatGPT 背后的相同生成式 AI 用于企业级项目。
参考文献
9 探索 GPTs
在 Discord 上加入我们的书籍社区

在之前的几个章节中,我们看到了许多如何利用 ChatGPT 进行各种活动的示例,从个人生产力到营销,从研究到软件开发。对于这些场景中的每一个,我们总是面临着类似的情况:我们从像 ChatGPT 这样的通用模型开始,然后提出非常具体的问题或提供特定的参考,以定制以满足我们的特定需求。
然而,如果我们的目标是为自己的用途获取一个极其专业化的模型,有时这可能是不够的。这就是为什么我们可能需要构建“特定用途的 ChatGPT”,而且幸运的是,OpenAI 本身开发了一个无代码平台来构建这些自定义助手,这些助手被称为 GPTs。
在本章中,我们将详细介绍 GPTs 的功能、能力和现实世界中的应用,涵盖了我们在之前的章节中看到的相同用例,以便你能够看到输出质量上的差异。此外,我们还将看到如何发布你的 GPT,并使其成为不仅适用于你自己,也适用于公司或所有人的生产应用。
在本章结束时,你将能够:
-
理解什么是 GPT 以及它能完成什么样的任务
-
在不编写任何代码的情况下构建你自己的 GPT
-
发布你的 GPT 并将其与外部系统集成
让我们从一些基本定义开始,然后进入实践。
技术要求
ChatGPT Plus
什么是 GPTs?
2023 年 11 月,OpenAI 引入了 GPTs,这是 ChatGPT 的专用版本,旨在提高生产力并满足特定的任务和需求。与通用 ChatGPT 不同,这些自定义版本被称为 GPTs,允许用户在没有任何任何编码知识的情况下创建定制的 AI 模型。
预先考虑分类法的考量是重要的,这在整个章节中都将相关。当我们提到 GPT 这个词时,你在书中已经发现并将会发现两种主要定义:
-
第一种是指 OpenAI 语言模型背后真正的“生成式预训练变换器”(
GPT)模型架构。我们已经在第一部分提到过这种架构,并且知道这是 ChatGPT 本身背后的框架。 -
第二种在更广义的情况下是指 OpenAI 允许用户以无代码方式创建的专用助手。通过 GPTs,OpenAI 指的是 ChatGPT 的专用版本,在这种语境下,单个 GPT 指你利用 GPTs 平台(所有 ChatGPT Plus 用户均可用)创建的一个助手。
在本章中,每当你读到 GPT 或 GPTs 时,请记住我们使用的是第二种定义。
GPTs 的理念就像 AI 代理(Agents)。事实上,通过 GPTs,我们正在构建由 LLM 驱动的实体,它们具有特定的指令,并提供了自定义知识库以及一组工具或插件来与周围环境交互。
让我们详细查看所有这些组件。首先,你可以全面查看所有已公发布的 GPTs。要做到这一点,你可以访问以下页面:chatgpt.com/gpts。

如从前面的图片所示,这是所有 GPTs 的市场。你可以通过类别(box 1)或通过说明你寻找的内容(box 2)来探索它。
然后要创建你自己的 GPT,你可以进入 OpenAI 账户并点击左上角的 Explore GPTs,然后点击 Create(如 box 3 所示)。

一旦进入编辑器,你将被要求配置你的 GPT,而在右侧你有机会进行实时测试。让我们探索所有组件:
-
Name(名称):你给你的 GPT 起的名字。
-
Description(描述):你的 GPT 功能的描述。如果你打算在市场发布你的 GPT,这非常重要,这样其他用户就可以轻松找到它(正如我们之前提到的,可以通过 Search GPTs 搜索栏通过自然语言搜索 GPTs)。
-
Instructions(指令):这是你 GPT 的元提示(metaprompt),即用自然语言编写的指令,用于根据你的特定需求定制助手,且最终用户看不到。
-
Conversation Starters(对话启动器):一组示例提示词,用户可以使用它们开始与 GPT 交互并建立信心。
-
Knowledge(知识):指的是我们可以为模型提供的自定义文档。当我们在这里上传文档时,我们的 GPT 将能够通过检索增强生成(RAG)模式遍历它,因此我们可以根据需要提供额外的知识,甚至将助手的回答限制在自定义知识库中。
-
Capabilities(能力):这些指的是我们可以为 GPT 提供的一组内置插件,无需编写单行代码。你可以从上面的图中看到开箱即用的三个插件:
-
Web browsing(网络浏览)用于搜索网络并获取最新信息。
-
DALL-E 3 用于生成插图。
-
Code Interpreter and Data Analysis(代码解释器和数据分析)用于在沙箱 Python 环境中执行代码并与分析文件交互。
-
-
Actions(动作):动作也可以被视为插件,但它们与能力不同,因为它们不是内置的,而是由 GPT 开发者指定的。例如,你可能会生成一个动作
在动作的背景下,如果你点击
Add Actions然后点击Get help from ActionsGPT,你可以获得原生集成在配置面板中的专用 GPT 的支持。


图 3:如何从 ActionsGPT 获取支持的示例
通过这种方式,你将被引导至你的 GPT 聊天界面:

图 4:ActionsGPT 的落地页。
看到我们正在见证“GPT 内部 GPT”的方法,这种感觉很神奇,你不觉得吗?
除了标准配置页外,还有另一种更具“对话性”的构建 GPT 的选项。事实上,你可以切换到 Create(创建)标签页,并用自然语言解释你想让 GPT 实现什么目标:

图 5:从自然语言对话开始创建 GPT 的示例。
我们将在本章的后续章节中看到这两种方法——标准配置和对话式配置。
既然我们已经知道了什么是 GPT,让我们来看看如何创建一个。在接下来的几个章节中,我们将创建 5 个不同的 GPT:
-
前四个将专注于我们已经用通用 ChatGPT 涵盖的四个领域——个人生产力、代码开发、营销和研究。其意图是比较特定领域 GPT 与通用 ChatGPT 的整体效率和准确性。
-
第五个将涵盖一个新领域,与艺术创作相关。我们将利用内置的
DALL-E插件以及由设计公司开发的其他插件(如Canvas)。
让我们从通过专门的 GPT 提升个人生产力开始。
个人助手
在这种场景下,我们将构建一个 GPT 来增强我们的健身训练。为此,我们将利用内置插件。此外,我们还将添加一些相关文档,以便助手基于特定的知识库。
我的目标是拥有一个助手,它可以根据我的健身目标、可用时间、性别、年龄、偏好等为我制定训练计划。我还希望助手能对其建议背后的原因提供清晰的解释。
为了实现这一切,我想确保我的助手能够:
-
向我提问设计我我最佳训练所需的特定问题
-
提供与其回复相关的相关信息和来源
-
考虑我的反馈,但如果它认为对我是正确的,它将能够坚持自己的观点
-
如果我的请求不合理或对健康有风险,则不满足我的请求
让我们按照上一节提到的所有配置步骤来看如何创建我们的 GPT:
-
Name (名称):我将我的助手命名为
WorkoutGPT。 -
Description (描述):这是我设置的描述:“帮助用户根据需求设计训练计划的健身助手。”
-
Instructions (指令):这里是我们 GPT 的真正核心。这是我提供 GPT 的指令集:
“你是一个健身 AI 助手,根据用户的需求帮助他们创建训练计划。
在生成计划之前,请确保询问以下问题:
- 健身目标和时间预期
- 年龄和性别
- 健身水平
- 训练的可用时间
- 为了制定合理的训练计划所需的其他所有元素(例如设备、潜在受伤情况……)
如有需要,使用提供的文档来丰富你的回答。
如果用户建议了某些不符合其目标,请坚持你的假设,并礼貌地解释背后的原因。
如果用户询问了可能对其健康产生风险的问题,请礼貌地建议放轻松并提供替代方案,同时解释背后的原因。”
-
Conversation Starters (对话启动器):在这里我设置了三个不同训练示例:
-
我想在 6 个月后参加马拉松。
-
生成一个 30 分钟的无设备 HIIT 训练计划。
-
生成一个仅需哑铃的 45 分钟举重训练计划。
-
-
Knowledge (知识):在这里我上传了标准的国家力量与调节协会 (
NSCA) 训练负荷表,这是一个帮助运动员和教练确定不同练习和训练课程适当负荷的工具。它看起来像:

图 6:NSCA 训练负荷表。来源:https://www.nsca.com/contentassets/61d813865e264c6e852cadfe247eae52/nsca_training_load_chart.pdf
我几乎忘记了一个关键步骤——插图!这看起来可能很表面,但为你的 GPT 设置图标会让它更有吸引力,特别是如果你计划将其发布给所有用户。幸运的是,我们在配置面板中直接集成了 DALL-E:

图 7:如何设置你的 GPT 图标的示例。
配置看起来是这样的:

图 8:WorkoutGPT 配置页面。
太棒了!现在让我们一些示例对话。
首先,让我们选择关于马拉松的对话启动器:

图 9:GPT 评估用户整体目标和健身水平的问题示例。
正如你所看到的,我们的 WorkoutGPT 会立即询问所需的信息。一旦提供了上述信息,我的助手生成了以下计划(此处输出被截断):

图 10:为马拉松训练生成的训练表示例。
此外,它如下指定了周五的力量训练:

图 11:GPT 生成的力量训练示例。
让我们关注力量训练。我想更好地了解如何调整重量。下面的插图显示了回复的第一部分:

图 12:GPT 解释如何在力量训练中确定重量的示例。
在同一个回复中,助手还引用了作为知识库提供的 NSCA 训练负荷表:

图 13:GPT 从自定义知识库检索信息的示例。
现在我想挑战我的 WorkoutGPT,请求一些可能对我健康有害的东西。例如,在一个月内在没有任何经验的情况下准备马拉松显然是一个糟糕的主意。一旦我提供了其启动问题的答案列表,让我们看看我的助手对此的看法:

图 14:考虑到与请求相关的风险,礼婉地引导用户改变目标和预期的示例。
正如你所看到的,我们的 WorkoutGPT 正在引导我们改变比赛方法。虽然它仍然为我们提供为期 3 周的跑步训练计划(此处输出被截断),但其目标不是在 3 小时 15 分钟内完赛马拉松,而是增强耐力和力量。
注意,如果我们向通用 ChatGPT 提问相同的问题,它的回复将如下:

图 15:ChatGPT 在存在相关风险的情况下仍完成用户任务的示例。
请注意,尽管 ChatGPT 口头上头上表达了顾虑,但它仍然在配合我的请求,为我提供了一份运行全程马拉松的计划。这可能会鼓励我——一个认为跑马拉松是个笑话的鲁莽初学者——投身于这场愚蠢的冒险,并对我的健康造成严重后果。
总的来说,GPTs 允许你非常具体地规定助手应该如何行为,以及它应该避免说什么或满足什么要求。
代码助手
在这一节中,我想开发一个专为数据科学项目定制的助手。更具体地说,我希望我的助手能够:
-
根据用户的任务,就如何设置数据科学实验提供清晰的指导
-
生成运行该实验所需的
Python代码 -
利用
code interpreter(代码解释器)能力来运行和检查代码 -
将最终代码推送到
GitHub仓库。
让我们逐步查看如何实现。
1. 设置指令
这个 GPT 是一个数据科学助手,帮助用户设置和运行数据科学实验。它提供了如何设置和运行数据科学实验的指导。它提供了如何组织实验的清晰指导,生成运行实验所需的 Python 代码,并利用 code interpreter 能力来运行和检查代码。GPT 将接收用户的输入并执行代码,并将最终代码推送到 GitHub 仓库。
2. 设置(可选)提示词
-
我如何设置实验?
-
生成
Python代码模型。 -
你能帮我设置吗?
3. 启用代码解释器与数据分析

图 16:启用代码解释器与数据分析。
4. 创建 GitHub 的 Action
为了与 GitHub 交互,我们需要
Schema(架构)。
定义架构的步骤
-
数据(Data):确定数据类型。
-
属性(Properties):
Schema架构下的字段。 -
必填字段(Required fields):
Schema架构中的必填项。 -
响应(Responses):
Schema架构的响应。
考虑以下示例:
openapi: 3.0
info:
title: ChatGPT API
description: GitHub action
version: 1.0
servers:
- url: https://api.github.com
description: GitHub server
paths:
/repos/{owner}/{repository}/contents/{path}:
put:
operationId: updateFile
summary: 在仓库中更新文件
description: 使用此端点创建或更新文件。
parameters:
- name: owner
in: path
required: true
description: 仓库的所有者。
schema:
type: string
- name: repository
in: path
required: true
description: 仓库名称。
schema:
type: string
- name: path
in: path
required: true
description: 仓库中的文件路径。
schema:
type: string
requestBody:
required: true
content:
application/json:
schema:
type: object
properties:
message:
type: string
description: 文件的提交消息。
content:
type: string
description: 新文件内容,Base64 编码。
sha:
type: string
description: 被替换文件的 SHA(如果是更新)。
branch:
type: string
description: 创建更新文件的分支。
committer:
type: object
properties:
name:
type: string
email:
type: string
required:
- message
- content
responses:
'200':
description: 文件更新或创建成功。
'201':
description: 文件创建成功。
'422':
description: 验证失败或文件已存在。
security:
- bearerAuth: []
components:
securitySchemes:
bearerAuth:
type: http
scheme: bearer
bearerFormat: token
这是 GitHub Action 的配置:

图 17:GitHub Action 的配置。
为了允许 Action 与我的仓库通信,我在 GitHub 个人设置(Settings > Developer settings > Personal access token)中创建了一个个人访问令牌。
然后,你可以按照如下方式测试连接:
太棒了!如你所见,我们现在有一个包含预定义内容的新文件。
现在让我们创建并测试它:
-
让我们先从一个关于如何设置分类实验的简单问题开始(截断输出):
![图 20:DataScience Assistant 提供如何处理分类问题的指导示例。]()
图 20:DataScience Assistant 提供如何处理分类问题的指导示例。
-
按照这些指令,我们现在准备好设置我们的实验了。我们想要处理著名的泰坦尼号乘客生还预测问题。我们将上传数据集(你可以在网上找到很多免费版本,我在
github.com/datasciencedojo/datasets/blob/master/titanic.csv下载了我的数据),并利用Logistic Regression模型。泰坦尼号生还预测任务是数据科学和机器学习中的一个经典问题。其目标是根据年龄、性别、客舱等级、票价等各种特征,预测泰坦尼号上的乘客是否会存活。这项通常用于教学分类技术,即在标记的数据集上训练模型,然后用于预测新数据的结果。挑战在于选择相关特征、处理缺失数据以及选择合适的机器学习算法以实现准确的预测。
-
这是我的查询:

图 21:DataScience Assistant 设计泰坦尼号生还实验的示例。
现在我将分享模型对每个步骤响应的某些截图:
-
加载数据:
![图 22:DataScience Assistant 执行实验第一步的示例。]()
图 22:DataScience Assistant 执行实验第一步的示例。
-
探索并清洗数据:

图 23:DataScience Assistant 探索和清洗数据的示例。
注意,当你看到符号 [>_] 时,意味着触发了代码解释器插件(code interpreter)。你可以点击它查看执行的代码:

图 24:DataScience Assistant 利用代码解释器插件生成的代码示例。
-
特征工程:
![图 25:DataScience Assistant 进行特征工程的示例。]()
图 25:DataScience Assistant 进行特征工程的示例。
-
数据划分:
![图 26:DataScience Assistant 将数据集划分为训练集和测试集的示例。]()
图 26:DataScience Assistant 将数据集划分为训练集和测试集的示例。
-
训练模型:
![图 27:DataScience Assistant 训练 Logistic Regression 模型示例。]()
图 27:DataScience Assistant 训练 Logistic Regression 模型示例。
-
评估模型:

图 28:DataScience Assistant 评估模型输出的示例。
这太酷了!它非常准确并且可以节省大量时间。此外,如果你考虑到在大型企业中处理许多项目的数据科学家,拥有一个类似的助手也有助于在项目之间遵循固定标准,从而使跨团队的维护保持一致。
我们要求 GPT 执行的最后一件事是将代码推送到我们的仓库(Repo)。让我们看看它是如何工作的:

图 29:DataScience Assistant 利用 Action 将代码推送到 GitHub 的示例。
如果我们点击该链接,就可以看到文件已成功上传:

图 30:通过 DataScience Assistant Action 上传的文件。
它成功了!再次,这是 GPTs 如何提高开发人员和数据科学家效率的示例。通过遵循通用框架设计数据科学实验并启用无需切换到 GitHub 即可推送的工作流,可以节省以前的时间,让数据科学家能够专注于项目的核心方面。
营销助手
正如我们在第 6 章——“使用 ChatGPT 精通营销”中所看到的,AI 助手在这一领域具有极高的价值。事实上,生成文本内容——如社交媒体帖子、博客文章或营销活动——可能是这些模型表现最佳的活动。
文案撰稿人员是专长撰写有说服力且吸引人的内容的专业,通常用于营销和广告目的。他们的工作通常包括编写促销材料,如广告、折页、网站、电子邮件、社交媒体帖子以及其他形式的内容,旨在说服观众采取特定行动,例如购买或订阅服务。
在本节中,我们将创建一个为此类活动定制的文案写作助手。为此,我将我的助手命名为 Copywrite Companion,并设置了以下配置组件:
-
指令(Instructions):
Copywrite Companion 是一个多功能助手,旨在帮助用户完成各种写作任务。它擅长根据用户输入生成产品表(包括文本和图像)、策划用于通讯或营销目的邮件活动、为新产品创建吸引人的视觉效果,以及为不同平台编写社交媒体帖子。它确保内容具有说服力、吸引力并符合目标受众,有效推广产品或服务。助手会根据需求调整风格,并在所有输出中追求创意、清晰和相关性。它以轻松、友好的语气进行交流,让互动感到亲切自然。 -
对话启动词(Conversation starters):
你能为我的新产品写描述吗?我需要为关于登山的文章起一个引人的标题给我几个关于跑步的博客文章的想法 -
能力(Capabilities):

图 31:Copywrite Companion 开启的插件。
让我们来看看实际效果:
-
我首先让它根据提供的图片编写一份产品表(输出截断):
![图 32:Copywrite Companion 生成产品描述的示例。]()
图 32:Copywrite Companion 生成产品描述的示例。
-
作为文案,我们可能想将这些信息插入到更有结构的存储库中,例如 Excel 文件。让我们要求助手这样做,利用它的代码解释器插件:
![图 33:Copywrite Companion 利用代码解释器插件将之前的响应转换为 文件的示例。]()
图 33:Copywrite Companion 利用代码解释器插件将之前的响应转换为 文件的示例。
-
这就是最终结果:

图 34:Copywrite Companion 生成的文件。
现在让我们让助手为我们的鞋子生成一条 Instagram 帖子:
如你所见,助手还提出了一个图像描述以利用 DALL-E 插件。由于这个描述在我看来合理,我将要求它生成图像:

图 36:Copywrite Companion 利用 DALL-E 插件生成图像的示例。
最后,我想了解一下竞争对手品牌——如 Nike 和 Adidas——是如何进行营销活动的。为了实现这一目标,我将要求我的助手从网络上收集一些证据:

图 37:Copywrite Companion 利用网页浏览插件进行竞争分析的示例。
如你所见,我们的助手正确地利用了网页浏览插件来检索所需的信息。此外,它还为我们提供了强大的洞察,让我们了解这两家竞争公司正在投入哪些杠杆,以便我们思考(或询问我们的助手)独特的差异化优势,让我们的品牌在竞争激的市场中脱颖而出。
我们还可以更进一步,利用代码解释器和数据分析插件获取有关竞争更具体的见解。例如,假设我们收集了一个具有以下结构的 Excel 表格:

图 38:Excel 表格上的竞争分析。
现在我们想从中生成一些可视化图表。让我们我们的 Copywriter Companion 来这样做吧:

图 39:Copywrite companion 生成柱状图和散点图的示例。

图 40:Copywrite companion 生成折线图的示例。
包括按照要求的执行报告:

图 41:Copywrite companion 生成执行报告的示例。
总的来说,针对营销活动对 ChatGPT 进行定制化,在生成新内容、设计营销策略以及在网络上进行竞争分析方面是非常有用的。
研究助手
在这一场景中,我们将再次关注研究,但这次将特别关注论文的检索。更具体地说,我们希望我们的助手能够执行以下操作:
-
从我们提供的自定义知识库中检索信息
-
将自定义文档与仅来自
Arxiv的论文相结合,并为此项任务启用网页插件 -
从
Database(在我们的情况下,它将托管在Notion中)检索其他研究人员正在进行的工作,这样我们就不会冒着撰写他人涵盖的文章的风险。
让我们来看看所有的步骤。
- 上传自定义文档。 出于目的,我将使用两篇关于机器学习中图像分类的论文:来自
J. Bird等人的“CIFAKE: AI 生成合成图像的图像分类与可解释性识别”,以及来自Khalis等人的“图像分类任务中视觉变换器(Vision Transformers)的综合研究”。
你可以在配置面板的相应部分上传这些文件:

图 42:在配置面板中上传自定义文档。
- 将自定义文档与网络引用整合。 为了实现这一点,我们需要启用网页插件:

图 43:在配置面板中启用网页浏览插件。
此外,我们还需要指定助手仅导航通过 Arxiv 存档。我们将在创建的指令集中查看如何指定该内容。
- 从 Notion 数据库检索信息。 这里的想法是,作为研究人员,我们可能会提出一些已经在由其他同事研究和开发的想法。假设我们在具有以下结构的
Notion Database(Not 数据库)中跟踪所有正在进行的研究:

图 44:Notion 数据库结构。
为了实现这一点,我们需要创建 GPT actions(GPT 操作)。为此需要遵循两个步骤:
-
在你的
Notion workspace中,你需要创建一个标记为内部(internal)的新连接,并创建一个新的Internal Integration Secret(或API key)。你可以将此连接命名为chatgpt或类似的名称。 -
在你的
GPT configuration pane中,你需要设置一个新的Action架构(schema)。由于在我们的情况下我们需要查询特定的Database,架构如下: -
涉及到身份验证时,你可以点击“Authentication”(认证)并选择“API Key”。输入以下信息:
-
API Key:使用来自Notion中新创建连接的Internal Integration Secret -
Auth Type:Bearer
-

图 45:Notion Action Schema。
整个架构的样貌如下:
openapi: 3.1.0
info:
title: Notion API
description: 与 Notion 页面、数据库和用户交互的 API。
version: 1.0.0
servers:
- url: https://api.notion.com/v1
description: Main API server
paths:
/databases/{database_id}/query:
post:
operationId: queryDatabase
summary: 查询数据库
parameters:
- name: database_id
in: path
required: true
schema:
type: string
- name: Notion-Version
in: header
required: true
schema:
type: string
example: 2022-06-28
constant: 2022-06-28
requestBody:
required: true
content:
application/json:
schema:
type: object
properties:
filter:
type: object
sorts:
type: array
items:
type: object
start_cursor:
type: string
page_size:
type: integer
responses:
'200':
description: 成功响应
content:
application/json
schema:
type: object
properties:
object:
type: string
results:
type: array
items:
type: object
next_cursor:
type: string
has_more:
type: boolean
太好了,现在我们已经准备好了所有的素材,我们需要设置系统消息(system message)以及可选的对话启动器。在这种场景中,我设置了以下指令:
你是一个 AI 研究助手,任务是利用各种工具和资源帮助研究人员。你的主要职责包括从自定义知识库、Arxiv 学术论文以及存储在 Notion 数据库中的正在进行的项目检索、整合和交叉引用信息。遵循这些指南以确保你的帮助准确、全面并避免重复:
-
始终基于提供的文档回答
-
如果你觉得需要,可以扩展网络搜索。你可以导航的唯一网站是 Arxiv。
-
如果用户询问,检查该主题是否已经在 Notion DB 中涵盖。
以及以下对话启动器:
-
VGGNet、ResNet 和 Inception 在图像分类的准确性和效率方面有何区别?
-
图像分类模型面临哪些挑战?数据增强和迁移学习等技术如何帮助它们?
3. CNN 深度如何影响图像分类?ResNet 的残差连接又是何帮助?
最终效果如下:

图 46:我们的 ResearchGPT 落地页
让我们测试一下:
-
我首先会问一个通用问题,并将根据提供的论文进行回答:
![图 47:ResearchGPT 从提供的文档检索知识的示例]()
图 47:ResearchGPT 从提供的文档检索知识的示例
-
接着,我想看看该主题是否已经被涵盖:
![图 48:ResearchGPT 使用预定义动作与 Notion 通信的示例]()
图 48:ResearchGPT 使用预定义动作与 Notion 通信的示例
-
最后,我想整合该主题使其使其独特性:
![图 49:ResearchGPT 使用浏览器插件的示例]()
图 49:ResearchGPT 使用浏览器插件的示例
如你所见,我们的助手可以利用我们提供的所有工具,并在需要时调用 Notion DB。
总结
在本章中,我们探索了如何获取定制的 GPTs 来实现你的特定目标。定制 ChatGPT 的可能性开启了全新的场景图景,高度专业的 AI 助手将成为专业人士的日常伙伴。此外,OpenAI 还提供了一个无代码界面来创建 GPTs,使得所有开发者不仅能够从扩展的可用解决方案市场中受益,还能从自己的创作中受益。
凭借插件扩展性和自定义知识库,自定义 GPTs 可以为你完成无限任务。然后,通过添加强大的动作(actions),它们还可以与周围环境通信,并从“仅仅”生成进化为自动化。
通过本章,我们也结束了本书的第二部分,该部分重点关注 ChatGPT 的实际应用。从下一章开始,我们将详细详细介绍大型企业如何利用 OpenAI 模型将其嵌入到业务流程中。我们将涵盖流行的企业用例,并展示如何通过 REST API 将 OpenAI 的大语言模型(LLMs)嵌入到应用程序中的实际实现方法。
参考文献


Practical Generative AI with ChatGPT, Second Edition: Start your generative AI journey by using ChatGPT to maximize your productivity and creativity
Welcome to Packt Early Access. We’re giving you an exclusive preview of this book before it goes on sale. It can take many months to write a book, but our authors have cutting-edge information to share with you today. Early Access gives you an insight into the latest developments by making chapter drafts available. The chapters may be a little rough around the edges right now, but our authors will update them over time.
You can dip in and out of this book or follow along from start to finish; Early Access is designed to be flexible. We hope you enjoy getting to know more about the process of writing a Packt book.
-
Chapter 1: Introduction to Generative AI
-
Chapter 2: Introducing OpenAI and ChatGPT
-
Chapter 3: From Prompt Design to Prompt Engineering
-
Chapter 4: Boosting Day-to-Day Productivity with ChatGPT
-
Chapter 5: Developing the Future with ChatGPT
-
Chapter 6: Mastering Marketing with ChatGPT
-
Chapter 7: Research Reinvented with ChatGPT
-
Chapter 8: Unleashing Creativity Visually with ChatGPT
-
Chapter 9: Exploring GPTs
-
Chapter 10: Leveraging OpenAI’s Models for Enterprise- Scale Applications with Models’ APIs
-
Chapter 11: Trending Use Cases
-
Chapter 12: Epilogue and Final Thoughts
1 Introduction to Generative AI
Join our book community on Discord

https://packt.link/EarlyAccess
Hello! Welcome to The Ultimate Guide to ChatGPT and OpenAI! In this book, we will explore the fascinating world of Generative Artificial Intelligence (GAI) and its groundbreaking applications. Generative AI has transformed the way we interact with machines, enabling computers to create, predict, and learn without explicit human instruction. Since the launch of ChatGPT in November 2023, we have witnessed unprecedented advances in natural language processing, image and video synthesis, and many other fields. Whether you are a curious beginner or an experienced practitioner, this guide will equip you with the knowledge and skills to navigate the exciting landscape of generative AI. So, let’s dive in and start with some definitions of the context we are moving in.
This chapter provides an overview of the field of generative AI, which consists of creating new and unique data or content using machine learning (ML) algorithms.
It focuses on the applications of generative AI to various fields, such as image synthesis, text generation, and music composition, highlighting the potential of generative AI to revolutionize various industries. This introduction to generative AI will provide context for where this technology lives, as well as the knowledge to collocate it within the wide world of AI, ML, and Deep Learning (DL). Then, we will dwell on the main areas of applications of generative AI with concrete examples and recent developments so that you can get familiar with the impact it may have on businesses and society in general.
Also, being aware of the research journey toward the current state of the art of generative AI will give you a better understanding of the foundations of recent developments and state-of-the-art models.
All this, we will cover with the following topics:
-
Understanding generative AI
-
Exploring the domains of generative AI
-
The history and current status of research on generative AI
By the end of this chapter, you will be familiar with the exciting world of generative AI, its applications, the research history behind it, and the current developments, which could have – and are currently having – a disruptive impact on businesses.
Introducing generative AI
AI has been making significant strides in recent years, and one of the areas that has seen considerable growth is generative AI. Generative AI is a subfield of AI and DL that focuses on generating new content, such as images, text, music, and video, by using algorithms and models that have been trained on existing data using ML techniques.
In order to better understand the relationship between AI, ML, DL, and generative AI, consider AI as the foundation, while ML, DL, and generative AI represent increasingly specialized and focused areas of study and application:
-
AI represents the broad field of creating systems that can perform tasks, showing human intelligence and ability and being able to interact with the ecosystem.
-
ML is a branch that focuses on creating algorithms and models that enable those systems to learn and improve themselves with time and training. ML models learn from existing data and automatically update their parameters as they grow.
-
DL is a sub-branch of ML, in the sense that it encompasses deep ML models. Those deep models are called neural networks and are particularly suitable in domains such as computervision or Natural Language Processing (NLP). When we talk about ML and DL models, we typically refer to discriminative models, whose aim is that of making predictions or inferencing patterns on top of data.
-
And finally, we get to generative AI, a further sub-branch of DL, which doesn’t use deep Artificial Neural Networks to cluster, classify, or make predictions on existing data: it uses those powerful Artificial Neural Network models to generate brand new content, from images to natural language, from music to video.
The following figure shows how these areas of research are related to each other:

Figure 1.1 – Relationship between AI, ML, DL, and generative AI
Generative AI models can be trained on vast amounts of data and then they can generate new examples from scratch using patterns in that data. This generative process is different from discriminative models, which are trained to predict the class or label of a given example.
The type of Neural Networks that feature Generative AI are called Large Foundation Models LFMs) and, in the case of language models like ChatGPT, we talk about Large Language Models (LLMs).
Large Language Models are a type of Artificial Neural Network featured by a transformer architecture. They are characterized by a huge number of parameters (in the order of trillions of parameters) and have been trained on billions of words. Given the training set, LLMs are capable of understanding and generating natural language given user’s input.
Even though text understanding and generation is probably one of the most outstanding features of Generative AI, this field covers many domains.
Domains of generative AI
In recent years, generative AI has made significant advancements and has expanded its applications to a wide range of domains, such as art, music, fashion, architecture, and many more. In some of them, it is indeed transforming the way we create, design, and understand the world around us. In others, it is improving and making existing processes and operations more efficient.
The fact that generative AI is used in many domains also implies that its models can deal with different kinds of data, from natural language to audio or images. Let us understand how generative AI models address different types of data and domains.
Text generation
One of the greatest applications of generative AI—and the one we are going to cover the most throughout this book—is its capability to produce new content in natural language. Indeed, generative AI algorithms can be used to generate new text, such as articles, poetry, and product descriptions.
For example, a language model such as GPT-4o, developed by OpenAI, can be trained on large amounts of text data and then used to generate new, coherent, and grammatically correct text in different languages (both in terms of input and output), as well as extracting relevant features from text such as keywords, topics, or full summaries.
Here is an example of working with GPT-3:

Figure 1.2 – Example of ChatGPT responding to a user prompt, also adding references
Next, we will move on to image generation.
Image generation
One of the earliest and most well-known examples of generative AI in image synthesis is the Generative Adversarial Network (GAN) architecture introduced in the 2014 paper by I. Goodfellow et al., Generative Adversarial Networks. The purpose of GANs is to generate realistic images that are indistinguishable from real images. This capability had several interesting business applications, such as generating synthetic datasets for training computer vision models, generating realistic product images, and generating realistic images for virtual reality and augmented reality applications.
Here is an example of faces of people who do not exist since they are entirely generated by AI:

Figure 1.3 – Imaginary faces generated by GAN StyleGAN2 at https://this-person-does-not-exist.com/en
Then, in 2021, a new generative AI model was introduced in this field by OpenAI, DALL-E. Different from GANs, the DALL-E model is designed to generate images from descriptions in natural language (GANs take a random noise vector as input) and can generate a wide range of images, which may not look realistic but still depict the desired concepts.
DALL-E has great potential in creative industries such as advertising, product design, and fashion, among others, to create unique and creative images.
Since its first release to date (December 2024), DALL-E has improved dramatically, as you can see in the following examples. Let’s see below an artistic creation authored by DALL-E at the dawn of its life:

Figure 1.4 – Images generated by DALL-E with a natural language prompt as input
Let’s now see what DALL-E3, the most recent version of the model at the time of writing this book, can produce (here I’m using Microsoft Image Creator powered by DALL-E3. You can try it at https://copilot.microsoft.com/images/create):

Figure 1.5 – Images generated by DALL-E3 with a natural language prompt as input
It’s impressive to see the level of improvement of this model in less than 18 months. Funny enough, we are just scraping the surface of the massive improvements occurring over the last months in the field of Generative AI.
Music generation
The first approaches to generative AI for music generation trace back to the 50s, with research in the field of algorithmic composition, a technique that uses algorithms to generate musical compositions. In fact, in 1957, Lejaren Hiller and Leonard Isaacson created the Illiac Suite for String Quartet (https://www.youtube.com/watch?v=n0njBFLQSk8), the first piece of music entirely composed by AI. Since then, the field of generative AI for music has been the subject of ongoing research for several decades. Among recent years’ developments, new architectures and frameworks have become widespread among the general public, such as the WaveNet architecture introduced by Google in 2016, which has been able to generate high-quality audio samples, or the Magenta project, also developed by Google, which uses Recurrent Neural Networks (RNNs) and other ML techniques to generate music and other forms of art. Then, in 2020, OpenAI also announced Jukebox, a neural network that generates music, with the possibility to customize the output in terms of musical and vocal style, genre, reference artist, and so on.
Those and other frameworks became the foundations of many AI composer assistants for music generation. An example is Flow Machines, developed by Sony CSL Research. This generative AI system was trained on a large database of musical pieces to create new music in a variety of styles. It was used by French composer Benoît Carré to compose an album called Hello World (https://www.helloworldalbum.net/), which features collaborations with several human musicians.
Here, you can see an example of a track generated entirely by Music Transformer, one of the models within the Magenta project:

Figure 1.6 – Music Transformer allows users to listen to musical performances generated by AI
Another incredible application of generative AI within the music domain is speech synthesis. It is indeed possible to find many AI tools that can create audio based on text inputs in the voices of well-known singers.
For example, if you have always wondered how your songs would sound if Kanye West performed them, well, you can now fulfill your dreams with tools such as FakeYou.com (https://fakeyou.com/), Deep Fake Text to Speech, or UberDuck.ai (https://uberduck.ai/).

Figure 1.7 – Text-to-speech synthesis with UberDuck.ai
I have to say, the result is really impressive. If you want to have fun, you can also try voices from your all your favourite cartoons as well, such as Winnie The Pooh...
Let’s go even further. What if we could generate a song from scratch, just asking the GenAI to do that for us in natural language? Well, today we can do that seamlessly and without any knowledge about music. Among the GenAI products that are rising in the musing market today, a great example is Suno, whose mission is “[...] is building a future where anyone can make great music. Whether you're a shower singer or a charting artist, we break barriers between you and the song you dream of making. No instrument needed, just imagination. From your mind to music.” (source: https://suno.com/about).

Can you believe that it became my summer 2024 hit? If you want to create your summer hit too, you can try it for free at https://suno.com/create.
Video generation
Generative AI for video generation shares a similar timeline of development with image generation. In fact, one of the key developments in the field of video generation has been the development of GANs. Thanks to their accuracy in producing realistic images, researchers have started to apply these techniques to video generation as well. One of the most notable examples of GAN-based video generation is DeepMind’s Motion to Video, which generated high-quality videos from a single image.
and a sequence of motions. Another great example is NVIDIA’s Video-to-Video Synthesis (Vid2Vid) DL-based framework, which uses GANs to synthesize high-quality videos from input videos.
The Vid2Vid system can generate temporally consistent videos, meaning that they maintain smooth and realistic motion over time. The technology can be used to perform a variety of video synthesis tasks, such as the following:
-
Converting videos from one domain into another (for example, converting a daytime video into a nighttime video or a sketch into a realistic image)
-
Modifying existing videos (for example, changing the style or appearance of objects in a video)
-
Creating new videos from static images (for example, animating a sequence of still images)
In September 2022, Meta’s researchers announced the general availability of Make-A-Video (https://makeavideo.studio/), a new AI system that allows users to convert their natural language prompts into video clips. Behind such technology, you can recognize many of the models we mentioned for other domains so far – language understanding for the prompt, image and motion generation with image generation, and background music made by AI composers.
Now, everything we’ve mentioned above pales in comparison with the latest text-to-video models. To name one, OpenAI announced in February 2024 a new text-to-video model called SORA, releasing some early experiments that were something unexpected:


Figure 1: Videos generated by SORA from a prompt in natural language. Source: https://openai.com/index/sora/
I do encourage you to visit the SORA webpage to have a look at the amazing videos it created. At the time of writing, SORA is not publicly available, as it is going through several tests ran by the OpenAI Red Team.
Overall, generative AI has impacted many domains for years, and some AI tools already consistently support artists, organizations, and general users. The future seems very promising; however, before jumping to the ultimate models available on the market today, we first need to have a deeper understanding of the roots of generative AI, its research history, and the recent developments that eventually lead to the current OpenAI models.
The history and current status of research
In previous sections, we had an overview of the most recent and cutting-edge technologies in the field of generative AI, all developed in recent years. However, the research in this field can be traced back decades ago.
We can mark the beginning of research in the field of generative AI in the 1960s, when Joseph Weizenbaum developed the chatbot ELIZA, one of the first examples of an NLP system. It was a simple rules-based interaction system aimed at entertaining users with responses based on text input, and it paved the way for further developments in both NLP and generative AI. However, we know that modern generative AI is a subfield of DL and, although the first Artificial Neural Networks (ANNs) were first introduced in the 1940s, researchers faced several challenges, including limited computing power and a lack of understanding of the biological basis of the brain. As a result, ANNs hadn’t gained much attention until the 1980s when, in addition to new hardware and neuroscience developments, the advent of the backpropagation algorithm facilitated the training phase of ANNs. Indeed, before the advent of backpropagation, training Neural Networks was difficult because it was not possible to efficiently calculate the gradient of the error with respect to the parameters or weights associated with each neuron, while backpropagation made it possible to automate the training process and enabled the application of ANNs.
Then, by the 2000s and 2010s, the advancement in computational capabilities, together with the huge amount of available data for training, yielded the possibility of making DL more practical and available to the general public, with a consequent boost in research.
In 2013, Kingma and Welling introduced a new model architecture in their paper Auto-Encoding Variational Bayes, called Variational Autoencoders (VAEs). VAEs are generative models that are based on the concept of variational inference. They provide a way of learning with a compact representation of data by encoding it into a lower-dimensional space called latent space (with the encoder component) and then decoding it back into the original data space (with the decoder component).
The key innovation of VAEs is the introduction of a probabilistic interpretation of the latent space. Instead of learning a deterministic mapping of the input to the latent space, the encoder maps the input to a probability distribution over the latent space. This allows VAEs to generate new samples by sampling from the latent space and decoding the samples into the input space.
For example, let’s say we want to train a VAE that can create new pictures of cats and dogs that look like they could be real.
To do this, the VAE first takes in a picture of a cat or a dog and compresses it down into a smaller set of numbers into the latent space, which represent the most important features of the picture. These numbers are called latent variables.
Then, the VAE takes these latent variables and uses them to create a new picture that looks like it could be a real cat or dog picture. This new picture may have some differences from the original pictures, but it should still look like it belongs in the same group of pictures.
The VAE gets better at creating realistic pictures over time by comparing its generated pictures to the real pictures and adjusting its latent variables to make the generated pictures look more like the real ones.
VAEs paved the way toward fast development within the field of generative AI. In fact, only 1 year later, GANs were introduced by Ian Goodfellow. Differently from VAEs architecture, whose main elements are the encoder and the decoder, GANs consist of two Neural Networks – a generator and a discriminator – which work against each other in a zero-sum game.
The generator creates fake data (in the case of images, it creates a new image) that is meant to look like real data (for example, an image of a cat). The discriminator takes in both real and fake data, and tries to distinguish between them – it’s the critic in our art forger example.
During training, the generator tries to create data that can fool the discriminator into thinking it’s real, while the discriminator tries to become better at distinguishing between real and fake data. The two parts are trained together in a process called adversarial training.
Over time, the generator gets better at creating fake data that looks like real data, while the discriminator gets better at distinguishing between real and fake data. Eventually, the generator becomes so good at creating fake data that even the discriminator can’t tell the difference between real and fake data.
Here is an example of human faces entirely generated by a GAN:

Figure 1.8 – Examples of photorealistic GAN-generated faces (taken from Progressive Growing of GANs for Improved Quality, Stability, and Variation, 2017: https://arxiv.org/pdf/1710.10196.pdf)
Both models – VAEs and GANs – are meant to generate brand new data that is indistinguishable from original samples, and their architecture has improved since their conception, side by side with the development of new models such as PixelCNNs, proposed by Van den Oord and his team, and WaveNet, developed by Google DeepMind, leading to advances in audio and speech generation.
Another great milestone was achieved in 2017 when a new architecture, called Transformer, was introduced by Google researchers in the paper, – Attention Is All You Need, was introduced in a paper by Google researchers. It was revolutionary in the field of language generation since it allowed for parallel processing while retaining memory about the context of language, outperforming the previous attempts of language models founded on RNNs or Long Short-Term Memory (LSTM) frameworks.
Transformers were indeed the foundations for massive language models called Bidirectional Encoder Representations from Transformers (BERT), introduced by Google in 2018, and they soon become the baseline in NLP experiments.
Transformers are also the foundations of all the Generative Pre-Trained (GPT) models introduced by OpenAI, including GPT-3, the model behind ChatGPT, as well as other Large Foundation Models.
Although there was a significant amount of research and achievements in those years, it was not until the second half of 2022 that the general attention of the public shifted toward the field of generative AI.
Not by chance, 2022 has been dubbed the year of generative AI. This was the year when powerful AI models and tools became widespread among the general public: diffusion-based image services (MidJourney, DALL-E 2, and Stable Diffusion), OpenAI’s ChatGPT, text-to-video (Make-a-Video and Imagen Video), and text-to-3D (DreamFusion, Magic3D, and Get3D) tools were all made available to individual users, sometimes also for free.
This had a disruptive impact for two main reasons:
-
Once generative AI models have been widespread to the public, every individual user or organization had the possibility to experiment with and appreciate its potential, even without being a data scientist or ML engineer.
-
The output of those new models and their embedded creativity were objectively stunning and often concerning. An urgent call for adaptation—both for individuals and governments—rose.
Henceforth, in the very near future, we will probably witness a spike in the adoption of AI systems for both individual usage and enterprise-level projects.
18 months later: main trends and innovations
Fast forward from November 2022 to today, we have witnessed a huge amount of innovation in the field of GenAI. Many of these innovations are linked to the brand-new models developed and released to the public, like the OpenAI’s GPT-4o and DALL-E3, but also Google Gemini, Meta Llama 3, Microsoft Phi3 and many others.
However, the most remarkable achievements probably lie in the way we interact with and build applications around those models. In this section, we are going to explore three main advancements that have marked the most popular reference architectures for GenAI-powered applications.
Retrieval Augmented Generation
One of the first limitations of ChatGPT and, generally speaking, of LLMs, was that of the knowledge base cutoff. In fact, the knowledge of LLMs is limited to the training set (made od public information from the internet) they have been trained on and, as long as this can be exhaustive, it’s not up to date. Plus, it is missing the proprietary knowledge base that might be relevant for us or our organization. For example, if you ask ChatGPT “what are my company’s policy for employees health insurance?”, the model won’t be able to answer since it has no access to this information.
To bypass this limitation, a new framework was designed to allow LLMs to navigate through customized documentations that we can provide. This framework is called Retrieval Augmented Generation (RAG).
The idea behind RAG is that of decoupling the LLM from the knowledge base we want to navigate through. To make it possible, the external knowledge base needs to be transformed into numerical vectors through a process called embedding and stored into a specialized Vector Database.
An embedding is a way of representing high-dimensional, non-numeric data, such as words or sentences, in a lower-dimensional space, such as vector. A text embedding can capture the semantic and syntactic features of the text, such as meaning, context, and similarity.
Each embedding is a vector of floating-point numbers, such that the distance between two embeddings in the vector space is correlated with semantic similarity between two inputs
in the original format.
For example, if two concepts are similar, then their vector representations should also be similar.

RAG is made of three phases:
-
Retrieval: given a user’s query and its corresponding vector, the most similar pieces of documents (those corresponding to the vectors that are closer to the user query’s vector) are retrieved and used as the base context for the LLM.
![]()
-
Augmentation: the retrieved context is enriched through additional instructions, rules, safety guardrails and similar practices that are typical of prompt engineering techniques (we will cover the topic of prompt engineering in Chapter 3).
![]()
-
Generation: based on the augmented context, the LLM generates the response to the user’s query.

The overall process looks as follows:

RAG combines the strengths of generative models and information retrieval systems to enhance the quality and relevance of generated content. Traditional generative models rely solely on their training data to produce responses, which can sometimes result in outdated or irrelevant information. RAG addresses this limitation by integrating external knowledge bases during the generation process.
For example, OpenAI's ChatGPT, when enhanced with RAG, can pull in current data from reliable sources to answer a query about recent events, which would be beyond the scope of its training cut-off. This synergy of retrieval and generation broadens the applicability of generative AI, making it more robust and useful in dynamic, information-rich environments.
Multimodality
In the first paragraph of this Chapter, we saw the various domains of Generative AI, ranging from text to images, from videos to music. Typically, Large Foundation Models tend to me domain-specific, as we saw for Large Language Models in the case of language understanding and generation, or DALL-E3 in case of image generation.
However, the recent advances in Generative AI have enabled the development of Large Multimodal Models (LMMs) that can process and generate different types of data, such as text, images, audio, and video.
LMMs share with “standard” Large Language Models (LLMs) the capability of generalization and adaptation typical of Large Foundation Models. However, LMMs are capable of processing heterogeneous data with the idea of mirroring the way humans interact with the surrounding ecosystem – that is, with all our senses.
A great example of multimodal model is OpenAI’s GPT-4o, which is able to interact with users via text, images and audio. Let’s see the following example:

As you can see, the model was able to analyze the image and reason over it. Let’s now go ahead and ask the model to generate an illustration:

The most interesting fact about LMMs is that they preserve their reasoning capabilities, making them suitable for complex reasoning in heterogeneous data contexts. Let’s consider this last example (showing only the first lines of the response):

As you may imagine, this opens a landscape of applications in various industries, and we are going to see some concrete examples in the upcoming chapters.
AI Agents
In previous sections, we uncovered how LFMs are great when it comes to resonate and generate content. However, they lack one ability, that is the one of taking actions and interact with the surrounding ecosystem that goes beyond the single user. For example, what if we want our LLM to be able not only to generate an amazing LinkedIn post, but also publish it on our page?
To overcome this limitation, AI agents emerge as key players. But what exactly are they? Agents can be seen as AI systems powered by Large Language Models that, given a user’s query, are able to interact with the surrounding ecosystem to the extent to which we allow them to. The perimeter of the ecosystem is delimited by the tools (or plug-ins) we provide the agents with (in our previous example, we might provide the agent with a LinkedIn plug-in, so that it is able to post the generated content)
Agents are made of the following ingredients:
-
An LLM which acts as the reasoning engine of the AI system
-
A system message which instructs the Agent to behave and think in a given way. For example, you can design an Agent as a teaching assistant for students with the following system message: “You are a teaching assistant. Given a student’s query, NEVER provide the final answer, but rather provide some hints to get there”.
-
A set of tools the Agent can leverage to interact with the surrounding ecosystem.
AI Agents are a perfect representation of the meaning of “LLM as reasoning engine of an application”. In fact, the beauty of Agents is that they can pick the best tool to use to accomplish a user’s request. For example, let’s say we have an AI agent to produce LinkedIn content, and we provide it with two tools: a LinkedIn plug-in and a web search plug-in (each one with a proper description of its functionality). Then, let’s explore the behaviour of the agent upon two different questions:
-
Generate a story about a little dog walking around the mountains🡪the Agent will generate the story without involving any plug-in.
-
Generate a story about the current weather in Milan🡪the Agent will invoke the web search plug-in to get the current weather in Milan.
-
Generate a LinkedIn post about the current weather in Mialn and publish it on my profile:the Agent will invoke the web search plug-in to get the current weather in Milan and the LinkedIn plug-in to post it on my profile.
The combinations of instructions and set of plug-ins make AI Agents extremely versatile, and you can create highly specialized entities to address specific scenarios.
And that’s not all.
Why having only one Agents if you can create your own crew of Agents talking to and cooperating among each other? Imagine multiple agents, each one with a specific expertise and goal, communicating and interacting among each other to accomplish a task. This is how multi-agent applications look like, and in the last few months, this pattern started showing emergent behaviors.
Let’s consider the following example. We want to generate an elevator pitch about the climate change. We need up to date information to do so (latest trends and research, future perspectives and so on), as well as solid research grounded by academic papers. Plus, we need to be concise yet sharp and effective, delivering all the key information in a very short pitch.
Now, we could ask all of that to a single agent, providing it with all the needed tools and with long instructions to accomplish the task. However, it is proven that LLMs tend to perform worse when we give them “too many things to do”. Instead, let’s use a multi-agent approach, creating a team with the following AI professionals:
-
A market analyst who can search the web for the latest news about climate change this will be an Agent with a web search plug-in and specific instructions to search for news;
-
An expert researcher who can easily navigate through academic research papers about climate change this will be an Agent with an Arxiv plug-in and specific instructions on how to retrieve relevant information;
-
A expert in public speaking who can easily consolidate all the information in one elevator pitch this will be an Agent with proper instructions on how to deliver perfect pitches;
-
A critic who will review the pitch and propose some changes to the expert in public speaking, if needed this will be an Agent with proper instructions on how to deliver perfect pitches;
So, upon the user’s question “Generate an elevator pitch about the current issue of climate change”, all the agents can start working on the project.
There are many frameworks that can help developer with multi-agent applications (including AutoGen, LangGraph, CrewAI), especially when it comes to the “flow” that we want our agents to follow. For example, we might want to enforce a specific number of iterations; or that all agents are invoked at least once; or even to involve us, as users, every iteration to provide further feedback to be incorporated in the upcoming iteration.
At the time of writing, multi-agent framework is showing promising advancements, yet it is still in an experimentation phase and probably still far from being enterprise-ready. Nevertheless, it is a glimpse of the outstanding reasoning capabilities behind LLMs and how they can unlock new ways of problem solving and creation.
Small Language Models
Large Language Models are, unsurprisingly, large. It means that the architecture of the Artificial Neural Network featuring LLMs is made of a huge amount of parameters, of the order of trillions. Ideally, the higher the number of parameters, the better the reasoning capabilities of the LLM. However, with high numbers of parameters come high cost of training and hosting, since a powerful AI infrastructure is needed. Plus, the energy consumption required by those models raise serious questions about the environmental impact of LLMs training and their overall sustainability in the long run. To give you an idea, these are some numbers related to open-source LLMs with different sizes (measured in number of parameters):

Figure 2: Carbon footprint and compute power required to train different models in the same datacenter. Source: https://arxiv.org/pdf/2302.13971
Luckily, over the last months, we witnessed a raising interest and investment in the research of smaller models, with the goal of maintaining decent reasoning capabilities while reducing the number of parameters. The idea is that smaller models, even if “less intelligent”, can still reach the desired requirements if specialized in more tailored activities. After all, do we really need huge and super powerful models of trillions of parameter to solve all the tasks that we have in mind?
These smaller models are called Small Language Models (SMLs) and, beside being lighter and less demanding in terms of infrastructure, they are also showing surprisingly high performance.
For example, if we consider the smaller version of Phi-3, part of the popular SLM Phi family developed by Microsoft, with “only” 7 billion parameters, we can see that it beats GPT-3.5-turbo in all the most popular LLMs benchmarks:

Figure 3: benchmark table of Phi-3 Small capabilities compared to other SLMs and LLMs. Source: https://azure.microsoft.com/en-us/blog/introducing-phi-3-redefining-whats-possible-with-slms/
Now, we might think that GPT-3.5-turbo is kind of “deprecated”, however we have to remember that, with its 175B parameters, it used to be the most powerful model in the market just 1 year ago, and it is remarkable to see that a 7B model is capable of better results.
That of SLMs is definitely a research stream to keep an eye on, especially when it comes to those scenarios where I might want to deploy my model locally or even customizing it with fine-tuning (we will cover fine tuning in the next chapter).
Summary
In this chapter, we explored the exciting world of generative AI and its various domains of application, including image generation, text generation, music generation, and video generation. We learned how generative AI models such as ChatGPT and DALL-E, trained by OpenAI, use DL techniques to learn patterns in large datasets and generate new content that is both novel and coherent. We also discussed the history of generative AI, its origins, and the current status of research on it.
The goal of this chapter was to provide a solid foundation in the basics of generative AI and to inspire you to explore this fascinating field further.
In the next chapter, we will focus on one of the most promising technologies available on the market today, ChatGPT: we will go through the research behind it and its development by OpenAI, the architecture of its model, and the main use cases it can address as of today.
References
-
This person does not exist: https://this-person-does-not-exist.com
-
https://azure.microsoft.com/en-us/blog/introducing-phi-3-redefining-whats-possible-with-slms/
2 OpenAI and ChatGPT: Beyond the Market Hype
Join our book community on Discord

https://packt.link/EarlyAccess
This chapter provides an overview of OpenAI and its most notable development—ChatGPT, highlighting its history, technology, and capabilities.
The overall goal is to provide a deeper knowledge of how ChatGPT can be used in various industries and applications to improve communication and automate processes and, finally, how those applications can impact the world of technology and beyond.
We will cover all this with the following topics:
-
What is OpenAI?
-
Overview of OpenAI model families
-
Road to ChatGPT: the math of the model behind it
-
ChatGPT: the state of the art
Technical requirements
In order to be able to test the example in this chapter, you will need the following:
-
An OpenAI account to access the Playground and the Models API (https://openai.com/api/login20)
-
Your favorite IDE environment, such as Jupyter or Visual Studio
-
Python 3.7.1+ installed (https://www.python.org/downloads)
-
pipinstalled (https://pip.pypa.io/en/stable/installation/) -
OpenAI Python library (https://pypi.org/project/openai/)
What is OpenAI?
OpenAI is a research organization founded in 2015 by Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, Wojciech Zaremba, and John Schulman. As stated on the OpenAI web page, its mission is “to ensure that Artificial General Intelligence (AGI) benefits all of humanity”. As it is general, AGI is intended to have the ability to learn and perform a wide range of tasks, without the need for task-specific programming.
Since 2015, OpenAI has focused its research on Deep Reinforcement Learning (DRL), a subset of machine learning (ML) that combines Reinforcement Learning (RL) with deep neural networks. The first contribution in that field traces back to 2016 when the company released OpenAI Gym, a toolkit for researchers to develop and test RL algorithms.

Figure 2.1 – Landing page of Gym documentation (https://www.gymlibrary.dev/)
OpenAI kept researching and contributing in that field, yet its most notable achievements are related to generative models—Generative Pre-trained Transformers (GPT).
After introducing the model architecture in their paper “Improving Language Understanding by Generative Pre-Training” and baptizing it GPT-1, OpenAI researchers soon released, in 2019, its successor, the GPT-2. This version of the GPT was trained on a corpus called WebText, which at the time contained slightly over 8 million documents with a total of 40 GB of text from URLs shared in Reddit submissions with at least 3 upvotes. It had 1.2 billion parameters, ten times as many as its predecessor.
Here, you can see the landing page of a UI of GPT-2 published by HuggingFace (https:// transformer.huggingface.co/doc/distil-gpt2):

Figure 2.2 – GPT-2 writing a paragraph based on a prompt. Source: https://transformer.huggingface.co/doc/distil-gpt2
Then, in 2020, OpenAI first announced and then released GPT-3, which, with its 175 billion parameters, dramatically improved benchmark results over GPT-2. It was with the GPT-3 model – more precisely, with its fine-tuned version called GPT-3.5 – that we entered the era of ChatGPT: in fact, at its first release, ChatGPT was powered by GPT-3.5.
From that moment (November 2022) to today, OpenAI released many new version of its GPT series: GPT-4, GPT-4-turbo, GPT-4vision (first multimodal model) and GPT-4o, where the “o” stands for “omni”, referring to its multimodal capabilities.
In addition to natural language generative models, OpenAI also developed in other fields of Generative AI, including:
-
DALL-E🡪 A text-to-image model that generates detailed and creative images from textual descriptions. It can produce a wide variety of images, from realistic scenes to imaginative and abstract art.
-
Whisper🡪 A speech recognition model designed to transcribe spoken language into text. It excels in understanding various accents, dialects, and noisy environments, making it highly effective for real-time transcription and voice-controlled applications.
-
SORA🡪A text-to-video model that generates realistic and imaginative videos from text instructions. SORA can create videos up to a minute long, maintaining visual quality and adherence to the user’s prompt. At the time of writing, SORA is not available to the general public yet.
Although OpenAI has invested in many fields of Generative AI, its contribution to text understanding and generation has been outstanding, thanks to the development of the foundation GPT models we are going to explore in the following paragraphs.
An overview of OpenAI Playground
Today, OpenAI offers a set of pre-trained, ready-to-use models that can be consumed by the general public. This has two important implications:
-
Powerful foundation models can be consumed without the need for long and expensive training
-
It’s not necessary to be a data scientist or an ML engineer to manipulate those models
Users can test OpenAI models in two ways:
-
OpenAI Playground, a friendly user interface where you can interact with models without the need to write any code.
-
Via REST API with a pro-code approach. In fact, all OpenAI models – and, generally speaking, LLMs – comes as pre-built components that can be consumed by applications via endpoints and keys.
In this chapter, we will explore the first approach, while we will see how to incorporate OpenAI models’ APIs in Chapter 11.
Before jumping into the Playground, let’s first have an overview of OpenAI’s models families.
OpenAI Models Families
Over the last years, OpenAI made huge advancements in the field of model developments, releasing newer models’ versions at insane speed. In this section, we will see the main models divided by domain.
-
Language Models🡪OpenAI's GPTs are advanced language models designed to generate human-like text based on given prompts. They are versatile and can be used for various natural language processing tasks such as text completion, translation, summarization, and coding. This is the field where OpenAI demonstrates outstanding performance thanks to its flagship models family: the GPTs. Since November 2022, when ChatGPT was launched, OpenAI released:
-
GPT-3.5-turbo, the model behind the first version of ChatGPT
-
GPT-4 (first multimodal model) and GPT-4-turbo (optimized for chats and assistants)
-
GPT-4o, the most advanced multimodal model
-
GPT-4o mini, the smaller version of the GPT-4o which is faster and cheaper.
-
In addition to those latest models, users can access a set of so called “GPT base models”, like the Babbage-002 or davinci-002 (the GPT-2), that represents the legacy completions APIs (we will cover completions APIs in the next section).
-
Image model🡪 OpenAI's image models, like DALL-E, are designed to generate and manipulate images from textual descriptions. DALL-E models can create highly detailed and imaginative visuals, enabling users to produce unique artwork, design concepts, and more. For instance, DALL-E 3 (the latest version of the model) is capable of creating intricate and creative images based on complex prompts, pushing the boundaries of what AI can do in the field of visual arts. These models are particularly useful in creative industries, digital marketing, and any field that benefits from high-quality, custom visuals.
-
Text to speech and Speech to text🡪 In Chapter 1, we already mentioned the Speech-to-text (STT) model developed by OpenAI, Whisper, that can handle and process audio input. In addition to that, OpenAI also developed Text-to-Speech (TTS) models that provide high-quality outputs that are both intelligible and expressive, making them suitable for a wide range of applications from customer service bots to educational tools. Text-to-Speech (TTS) models developed by OpenAI convert written text into spoken language, producing realistic and natural-sounding audio. These models are essential for creating voice assistants, improving accessibility for visually impaired users, and generating automated announcements.
-
Text to video🡪with the announcement of SORA, OpenAI revealed its cutting-edge text-to-video model capable of generating realistic and imaginative video scenes from text descriptions. By adapting techniques from DALL-E and incorporating transformers, SORA can create high-fidelity videos up to one minute long. It excels in maintaining 3D consistency, object permanence, and simulating interactions within videos. While still facing challenges like accurately modeling complex physics and ensuring temporal coherence, SORA holds significant potential for creative industries, offering new possibilities in video production and storytelling
-
Embeddings🡪Embeddings models from OpenAI transform text into numerical representations called vectors (or embeddings) that capture semantic meaning and are projected in a multi-dimensional vector space.
The mathematical distances between different instances in this space represent their similarity in terms of meaning. As an example, imagine the words queen, woman, king, and man. Ideally, in our multidimensional space, where words are vectors, if the representation is correct, we want to achieve the following:

Figure 2.10 – Example of vectorial equations among words
This means that the distance between woman and man should be equal to the distance between Queen and King.
Embeddings can be extremely useful in intelligent search scenarios. Indeed, by getting the embedding of the user input and the documents the user wants to search, it is possible to compute distance metrics (namely, cosine similarity) between the input and the documents. By doing so, we can retrieve the documents that are closer, in mathematical distance terms, to the user input.
The OpenAI’s embedding models text-embedding-ada-002 and text-embedding-3-large models offer state-of-the-art performance in tasks such as text similarity, text search, and code search by creating dense vector representations of text. These embeddings allow for efficient and effective comparison of large text datasets, improving search accuracy and relevance
- Moderation🡪OpenAI's moderation models are designed to detect and filter out inappropriate, harmful, or unsafe content in text. These models are integral to maintaining safe and respectful online environments by identifying potentially offensive or harmful language. The latest moderation model, text-moderation-007, is robust and effective in ensuring content compliance and safety across platforms. It helps developers and companies enforce community guidelines and prevent the spread of harmful content, thereby promoting safer digital interactions.
Overall, OpenAI has reached more and more coverage over the last years, addressing diverse areas of Generative AI with state-of-the-art models.
Trying OpenAI models in the Playground
To access your OpenAI playground, you need to create an OpenAI account and navigate through https://platform.openai.com/playground. This is how the landing page looks like:

Figure 2.4 – OpenAI Playground at https://platform.openai.com/playground
As you can see from Figure 2.4, the Playground offers a UI where the user can start interacting with the model, which you can select on the right-hand side of the UI.
Before diving deeper into the main sections of the playground, let’s first define some jargon you will see in this chapter:
-
Tokens: Tokens can be considered as word fragments or segments that are used by the API to process input prompts. Unlike complete words, tokens may contain trailing spaces or even partial sub-words. To better understand the concept of tokens in terms of length, there are some general guidelines to keep in mind. For instance, one token in English is approximately equivalent to four characters, or three-quarters of a word.
-
Prompt: In the context of natural language processing (NLP) and ML, a prompt refers to a piece of text that is given as input to an AI language model to generate a response or output. The prompt can be a question, a statement, or a sentence, and it is used to provide context and direction to the language model.
-
Context: In the field of GPT, context refers to the words and sentences that come before the user’s prompt. This context is used by the language model to generate the most probable next word or phrase, based on the patterns and relationships found in the training data.
-
Model confidence: Model confidence refers to the level of certainty or probability that an AI model assigns to a particular prediction or output. In the context of NLP, model confidence is often used to indicate how confident the AI model is in the correctness or relevance of its generated response to a given input prompt.
-
Functions. Functions is yet another definition for the concept of tools or plug-ins introduced in Chapter 1. With functions, we provide the model with an extra skill that it can invoke to accomplish user’s task. A function will always have a description in natural language, so that the model knows when to invoke it.
The preceding definitions will be pivotal in understanding how to use Azure OpenAI model families and how to configure their parameters.
In the Playground, there are four main sections to interact with the models:
Chat: Here you can test all the chat models available today, including both text-only models (like the GPT-3.5) or multimodal models (like the GPT-4o). You can provide a system message – that is the set of instructions that you provide your model with, all in natural language. You can also compare the output of two different models, given the same user’s question. The following is an example on how to do that:

Figure 2.5 – An example of comparison between two models.
For each model, you can also play with some parameters that you can configure. Here is a list:
-
Temperature (ranging from 0 to 1): This controls the randomness of the model’s response. A low-level temperature makes your model more deterministic, meaning that it will tend to give the same output to the same question. For example, if I ask my model multiple times What is OpenAI? with temperature set as 0, it will always give the same answer. On the other hand, if I do the same with a model with temperature set as 1, it will try to modify its answers each time, in terms of wording and style.
-
Max tokens: This controls the length (in terms of tokens) of the model’s response to the user’s prompt.
-
Stop sequences (user input): This makes responses end at the desired point, such as the end of a sentence or list.
-
Top probabilities (ranging from 0 to 1): This controls which tokens the model will consider when generating a response. Setting this to 0.9 will consider the top 90% most likely of all possible tokens. One could ask Why not set top probabilities as 1 so that all the most likely tokens are chosen? The answer is that users might still want to maintain variety when the model has low confidence, even in the highest-scoring tokens.
-
Frequency penalty (ranging from 0 to 1): This controls the repetition of the same tokens in the generated response. The higher the penalty, the lower the probability of seeing the same tokens more than once in the same response. The penalty reduces the chance proportionally, based on how often a token has appeared in the text so far (this is the key difference from the following parameter).
-
Presence penalty (ranging from 0 to 2): This is similar to the previous one but stricter. It reduces the chance of repeating any token that has appeared in the text at all so far. As it is stricter than the frequency penalty, the presence penalty also increases the likelihood of introducing new topics in a response.
Besides trying OpenAI models in the Playground, you can always call the models API in your custom code and embed models into your applications. Indeed, in the right corner of the Playground, you can click on View code and export the configuration as shown here:
- Assistants. OpenAI Assistants can be seen as a way to develop AI Agents (introduced in Chapter 1) faster and more easily. In fact, Assistants can be defined as entities powered by an LLM, with a set of instructions to follows and a set of tools or plug-ins to use – basically the same definition of AI Agents!
In the case of OpenAI assistants, they come with three pre-built tools:
-
File Search🡪it allows the user to upload custom documents so that the Assistant can navigate through it to accomplish the query. It operates with a RAG-based framework.
-
Function Calling🡪it allows the user to define a set of custom functions that can be invoked by the Assistant to accomplish a given task.
-
Code Interpreter🡪it refers to the capability of the Assistant to run code either against provided documents (for example, in case of spreadsheets or analytical papers that required mathematical computations) or simply to solve complex tasks provided by the user (for example, complex mathematical problems).
In the following picture, you can see an example of an Assistant called “Chat with PDF”, which is specialized in responding upon provided documents (in my case, I uploaded the paper “LLaMA: Open and Efficient Foundation Language Models” by Hugo Touvron et al.

Figure 2.8 – Example of an OpenAI Assistant.
As you can see from the preceding screenshot, the Assistant was able to answer my question retrieving knowledge from the provided document. In fact, my question was pretty vague, since the term “toxicity” can refer to multiple domains; nevertheless, the Assistant knows to watch over the provided documents as primary source of information.
Completions. This section refers to a class of models called “base models”, like the GPT-3, they are the basis from where the so called “assistant models” (or chat models, as we saw previously) are built. For example, the chat model GPT-3.5-turbo (the model behind ChatGPT), is a fine-tuned version of the base model GPT-3.
Completions (base) models are designed for generating single responses to prompts, making them suitable for tasks like text generation and summarization without maintaining context over multiple interactions. Chat (assistant) models, on the other hand, are optimized for interactive conversations, capable of maintaining context across multiple turns, and are ideal for applications like chatbots and virtual assistants.
Below you can see an example of a typical completion task in the Playground:

Figure 1: Example of completion task in OpenAI Playground
As you can see, upon my words “Today I went to a grocery store and”, the model completed the sentence with the most likely words.
Today, completion models are rarely used as they are outperformed by chat models, yet they can be further fine-tuned to tailored use cases (we will cover fine-tuning later on in this section).
- Text to speech. In addition to Whisper, the aforementioned speech-to-text model, OpenAI also released its TTS (text-to-speech) model that can be tested directly in the Playground.
Let’s see an example:

Figure 2: Example of using OpenAI’s TTS models in the Playground
As you can see from the above picture, you can select the voice, model, speed and format of the generated audio.
All the previous models come as pre-built, in the sense that they have already been pre-trained on a huge knowledge base.
However, there are some ways you can make your model more customized and tailored for your use case.
The first method is embedded in the way the model is designed, and it involves providing your model with the context in the few-learning approach (we will focus on this technique later on in the book). Namely, you could ask the model to generate an article whose template and lexicon recall another one you have already written. For this, you can provide the model with your query of generating an article and the former article as a reference or context, so that the model is better prepared for your request.
Here is an example of it:

Figure 2.11 – An example of a conversation within the OpenAI Playground with the few-shot learning approach
In the previous illustration, I instructed the model to output only the label of the tweet’s sentiment, providing it with three examples on how to do that.
The second method is more sophisticated and is called fine-tuning. Fine-tuning is the process of adapting a pre-trained model to a new task.
In fine-tuning, the parameters of the pre-trained model are altered, either by adjusting the existing parameters or by adding new parameters, to better fit the data for the new task. This is done by training the model on a smaller labeled dataset that is specific to the new task. The key idea behind fine-tuning is to leverage the knowledge learned from the pre-trained model and fine-tune it to the new task, rather than training a model from scratch. Have a look at the following figure:

Figure 2.12 – Model fine-tuning
In the preceding figure, you can see a schema on how fine-tuning works on OpenAI pre-built models. The idea is that you have available a pre-trained model with general-purpose weights or parameters. Then, you feed your model with custom data, typically in the form of key-value prompts and completions as shown here:
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
...
Once the training is done, you will have a customized model that performs particularly well for a given task, for example, the classification of your company’s documentation.
The nice thing about fine-tuning is that you can make pre-built models tailored to your use cases, without the need to re-train them from scratch, yet leveraging smaller training datasets and hence less training time and computing. At the same time, the model keeps its generative power and accuracy learned via the original training, the one that occurred on the massive dataset.
In this paragraph, we got an overview of the models offered by OpenAI to the general public, from those you can try directly in the Playground (GPT, Codex) to more complex models such as embeddings. We also learned that, besides using models in their pre-built state, you can also customize them via fine-tuning, providing a set of examples to learn from.
In the following sections, we are going to focus on the background of those amazing models, starting from the math behind them and then getting to the great discoveries that made ChatGPT possible.
ChatGPT: the state of the art
In November 2022, OpenAI announced the web preview of its conversational AI system, ChatGPT, available to the general public. This was the start of huge hype coming from subject matter experts, organizations, and general users – to the point that, after only 5 days, the service reached 1 million users!
Before writing about ChatGPT, I will let it introduce itself, using a snapshot taken a few days after the launch:

Figure 2.25 – ChatGPT introduces itself in November 2022.
As mentioned above, the first release of ChatGPT was built on top of an advanced language model that utilizes a modified version of GPT-3, which has been fine-tuned specifically for dialogue. This fine-tuned version is called GPT-3.5-turbo. The optimization process involved Reinforcement Learning with Human Feedback (RLHF), a technique that leverages human input to train the model to exhibit desirable conversational behaviors.
We can define RLHF as a machine learning approach where an algorithm learns to perform a task by receiving feedback from a human. The algorithm is trained to make decisions that maximize a reward signal provided by the human, and the human provides additional feedback to improve the algorithm’s performance. This approach is useful when the task is too complex for traditional programming or when the desired outcome is difficult to specify in advance.
The relevant differentiator here is that ChatGPT has been trained with humans in the loop so that it is aligned with its users. By incorporating RLHF, ChatGPT has been designed to better understand and respond to human language in a natural and engaging way.
The same RLHF mechanism was used for what we can think of as the predecessor of ChatGPT— InstructGPT. In the related paper published by OpenAI’s researchers in January 2022, InstructGPT is introduced as a class of models that is better than GPT-3 at following English instructions.
ChatGPT is still available today as a free application anyone can use; however, since February 2023, OpenAI announced a new paid version called ChatGPT Plus and costs 20$/month, offering subscribers several advantages, including access to the latest models, fastest response time, a great set of plug-ins and the possibility to create your own Assistant in the GPTs playground.
Let’s have a quick tour of ChatGPT user interface at the current date:

Figure 3: Landing page of ChatGPT at chatgpt.com
Let’s double click on each section:
-
You can decide the model to use behind ChatGPT. In my case, I have set the GPT-4o, which is only available for paid subscription. This is the model we are going to use throughout this book.
-
A set of pre-built prompts are proposed to the user to familiarize themselves with the application.
-
The text box is the place where users can ask their questions. Note that there is a small clip in the left-hand corner: it indicates the possibility of uploading files the model will be able to navigate through. Those files can be either uploaded locally, or retrieved from cloud storage like Google Drive or OneDrive.
-
On the left-hand sidebar, you can see the list of previous chats (mine are covered) you had with ChatGPT. This is an extremely useful tool, since in each chat, through the various round of interactions you had with the model, you created a context ChatGPT is aware of. This means that, if you want to continue a previously started conversation, you can open the related chat and start talking with the model without describing again the whole scenario.
-
A nice recently added feature is the possibility of having shared workspaces with other users, so that you can cooperate with your team members. Note that this is an additional paid feature which is not included in the Plus plan, and it costs extra 5$/month.
-
Finally, ChatGPT Plus offers the possibility of creating the GPTs, that are personalized assistant you can tailor for specific functions. You can decide to keep your GPT private or to publish it into the GPTs store, where anyone can use and rate it. We are going to cover the GPTs in a dedicated Chapter later on in this book.
Finally, ChatGPT Plus comes with two built-in plugins that the model can invoke autonomously if needed: the web search plug-in (to retrieve up-to-date information) and the code-interpreter plug-in (to process analytical documentation or perform complex mathematical tasks).
Throughout this book, we will leverage ChatGPT Plus to showcase the capabilities of the latest models and features, nevertheless the majority of examples that we will cover can be achieved with the free version of ChatGPT as well (which is currently powered by GPT-3.5-turbo).
The ongoing developments and improvements in ChatGPT’s architecture and training methods promise to push the boundaries of language processing even further.
Summary
In this chapter, we went through the history of OpenAI, its research fields, and the latest developments, up to ChatGPT. We went deeper into the OpenAI Playground for the test environment and how we can leverage all the OpenAI’s models families.
With this first glance at the OpenAI models, we will also saw how easy it is to test or embed pre-trained models into your applications: the game-changer element here is that you don’t need powerful hardware and hours of time to train your models, since they are already available to you and can also be customized if needed, with a few examples.
In the next chapter, we also begin Part 2 of this book, where we will see ChatGPT in action within various domains and how to unlock its potential. You will learn how to get the highest value from ChatGPT by properly designing your prompts, how to boost your daily productivity, and how it can be a great project assistant for developers, marketers, and researchers.
References
-
Radford, A., & Narasimhan, K. (2018). Improving language understanding by generative pre-training.
-
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. ArXiv. https://doi.org/10.48550/ arXiv.1706.03762 OpenAI. Fine-Tuning Guide. OpenAI platform documentation. https://platform.openai.com/docs/guides/fine-tuning.
3 Understanding Prompt Design
Join our book community on Discord

https://packt.link/EarlyAccess
In the previous chapters, we mentioned the term prompt several times while referring to user input in ChatGPT and LLMs models in general.
Since prompts have a massive impact on LLMs performance, prompt engineering is a crucial activity while designing LLM-powered applications. In fact, there are several techniques that can be implemented to not only to refine your LLM’s responses, but also to reduce risks associated with hallucination and biases.
In this chapter, we are going to cover the emerging techniques in the field of prompt engineering, starting from basic approaches up to advanced frameworks. More specifically, we will go through the following topics:
-
Introduction to prompt engineering
-
Zero, one and few shot learning
-
Basic principles of prompt engineering
-
Advanced techniques of prompt engineering
-
Mitigation of risks and hallucinations with prompt engineering
-
Handling prompt injections
By the end of this chapter, you will have the foundations to build functional and solid prompts for your LLM-powered applications, which will also be relevant in the upcoming chapters.
What is prompt engineering?
Before explaining what prompt engineering is, let’s start with the atomic definition of prompt.
A prompt is a text input that guides the behaviour of an LLM to generate a text output. In the context of LLMs and LLMs-powered applications, we can distinguish two types of prompts:
- The prompt that the user’s type when interacting with the LLM. For example, a prompt might be “Give me a detailed explanation of a proton”, or “Generate a workout plan to run a marathon”. Below you can see an example of a simple user prompt:


Figure 1: Example of user’s prompt.
You will hear referring to this component simply as “prompt”, “query”, or “user’s input”.
- The prompt that instruct the model to behave in a certain way regardless of the user’s query. This refers to the set of instructions in natural language that we provide the model with, so that it behaves in a certain way when interacting with end users. You can think about that as a sort of “backend” of your LLM, something that will be handled by the application developers rather than the final users.
In the following picture, you can see an example where we ask the model the same above questions, yet this time we add a system message to instruct the model to behave in a certain way:



Figure 2: Example of a metaprompt.
We refer to this type of prompt as “metaprompt” or “system message”.
Henceforth, prompt engineering is the process of designing effective prompts that elicit high-quality and relevant outputs from LLMs. Prompt engineering requires creativity, understanding of the LLM, and precision.

Figure 3: Example of prompts engineering to specialize LLMs.
Over the last years, prompt engineering became a brand new discipline itself, and this is a demonstration of the fact that, interacting with those models require a new set of skills and capabilities that did not exist before. The “art of prompting” has become a top skill when it comes to build Generative AI applications in enterprise scenarios; however, it can also be extremely useful for individual users that use ChatGPT or similar AI assistant in daily tasks, as it improves dramatically the quality and accuracy of results.
Over the next sections, we are going to see some examples of how to build efficient and robust prompts and metaprompts for your LLM applications. Note that all the techniques we are going to cover, can be embedded both in the user prompt and the metaprompt, depending on your specific requirements. To do that, I will leverage the OpenAI Playground, so that I can have greater flexibility to operate at both levels of prompting.
Zero-, one-, and few-shot learning – typical of transformers models
In the previous chapters, we mentioned how LLMs typically come in a pre-trained format. They have been trained on a huge amount of data and have had their (billions of) parameters configured accordingly.
However, this doesn’t mean that those models can’t learn anymore. In Chapter 2, we saw that one way to customize an OpenAI model and make it more capable of addressing specific tasks is by fine-tuning.
Fine-tuning is the process of adapting a pre-trained model to a new task. In fine-tuning, the parameters of the pre-trained model are altered, either by adjusting the existing parameters or by adding new parameters so that they fit the data for the new task. This is done by training the model on a smaller labeled dataset that is specific to the new task. The key idea behind fine-tuning is to leverage the knowledge learned from the pre-trained model and fine-tune it to the new task, rather than training a model from scratch.
Fine-tuning is a proper training process that requires a training dataset, compute power, and some training time (depending on the amount of data and compute instances).
That is why it is worth testing another method for our model to become more skilled in specific tasks: shot learning: The idea is to let the model learn from simple examples rather than the entire dataset. Those examples are samples of the way we would like the model to respond so that the model not only learns the content but also the format, style, and taxonomy to use in its response.
Furthermore, shot learning occurs directly via the prompt (as we will see in the following scenarios), so the whole experience is less time-consuming and easier to perform.
The number of examples provided determines the level of shot learning we are referring to. In other words, we refer to zero-shot if no example is provided, one-shot if one example is provided, and few-shot if more than 2-3 examples are provided.
Let’s focus on each of those scenarios:
-
Zero-shot learning. In this type of learning, the model is asked to perform a task for which it has not seen any training examples. The model must rely on prior knowledge or general information about the task to complete the task. For example, a zero-shot learning approach could be that of asking the model to generate a description, as defined in my prompt:
![Figure 4.3 – Example of zero-shot learning]()
Figure 4.3 – Example of zero-shot learning
-
One-shot learning: In this type of learning, the model is given a single example of each new task it is asked to perform. The model must use its prior knowledge to generalize from this single example to perform the task. If we consider the preceding example, I could provide my model with a prompt-completion example before asking it to generate a new one:

Figure 4.4 – Example of one-shot learning
{"prompt": "<prompt text>", "completion": "<ideal generated text>"}
Note that the way I provided an example was similar to the structure used for fine-tuning:
- Few-shot learning: In this type of learning, the model is given a small number of examples (typically between 2 and 5) of each new task it is asked to perform. The model must use its prior knowledge to generalize from these examples to perform the task. Let’s continue with our example and provide the model with further examples:

Figure 4.5 – Example of few-shot learning with three examples
The nice thing about few-shot learning is that you can also control model output in terms of how it is presented. You can also provide your model with a template of the way you would like your output to look. For example, consider the following tweet classifier:

Figure 4.6 – Few-shot learning for a tweets classifier. This is a modified version of the original script from https://learn.microsoft.com/en-us/azure/cognitive-services/openai/how-to/completions
Let’s examine the preceding figure. First, I provided ChatGPT with some examples of labeled tweets. Then, I provided the same tweets but in a different data format (list format), as well as the labels in the same format. Finally, in list format, I provided unlabeled tweets so that the model returns a list of labels.
Shot-learning possibilities are limitless– it’s only a matter of testing and a little bit of patience in finding the proper prompt design.
As mentioned previously, it is important to remember that these forms of learning are different from traditional supervised learning, as well as fine-tuning. In few-shot learning, the goal is to enable the model to learn from very few examples, and to generalize from those examples to new tasks.
Now that we’ve learned how to let OpenAI models learn from examples, let’s focus on how to properly define our prompt to make the model’s response as accurate as possible.
Principles of prompt engineering
Traditionally, in the context of computing and data processing, we often use the expression “garbage in, garbage out”, meaning that the quality of output is determined by the quality of the input. If incorrect or poor-quality data (garbage) is entered into a system, the output will also be flawed or nonsensical (garbage).
When it comes to prompting, the story is similar: if we want accurate and relevant results from our LLMs, we need to provide high-quality input. However, building good prompts is not just about the quality of the response. In fact, we can construct good prompts to:
-
Maximize the relevancy of LLM’s responses
-
Specify the type formatting and style of responses
-
Provide conversational context
-
Reduce inner biases and improve fairness and inclusivity
-
Reduce hallucination
In the context of large language models (LLMs), "hallucination" refers to the generation of text or responses that are factually incorrect, nonsensical, or not grounded in the training data. This occurs when an LLM produces confident-sounding but erroneous or fabricated information. Hallucinations can arise due to the probabilistic nature of these models, which may predict the next word or phrase based on patterns rather than verified facts.
Let’s see some basic techniques to achieve these results.
Clear Instructions
The principle of giving clear instructions is to provide the model with enough information and guidance to perform the task correctly and efficiently. Clear instructions should include the following elements:
-
The goal or objective of the task, such as “write a poem” or “summarize an article”.
-
The format or structure of the expected output, such as “use four lines with rhyming words” or “use bullet points with no more than 10 words each”.
-
The constraints or limitations of the task, such as “do not use any profanity” or “do not copy any text from the source”.
-
The context or background of the task, such as “the poem is about autumn” or “the article is from a scientific journal”.
Let’s say, for example, that we want our model to fetch any kind of instructions from text and return to us a tutorial in a bullet list. Plus, if there are no instructions in the provided text, the model should inform us about that.
To try so, let’s leverage the OpenAI Playground under the chat section, so that we can provide both the metaprompt and the prompt.
For this scenario, I’ll set the following metaprompt:
You are an AI assistant that helps human by generating tutorials given a text.
You will be provided with a text. If the text contains any kind of istructions on how to proceed with something, generate a tutorial in a bullet list.
Otherwise, inform the user that the text does not contain any instructions.
And this will be our user prompt:
To prepare the known sauce from Genova, Italy, you can start by toasting the pine nuts to then coarsely
chop them in a kitchen mortar together with basil and garlic. Then, add half of the oil in the kitchen mortar and season with salt and pepper.
Finally, transfer the pesto to a bowl and stir in the grated Parmesan cheese.
Let’s see how it works:

Figure 4: Example of clear instructions in the metaprompt.
Note that, if we pass the model another text which does not contain any instructions, it will be able to respond as we instructed it:

Figure 5: Example of chat model following instructions.
By giving clear instructions, you can help the model understand what you want it to do and how you want it to do it. This can improve the quality and relevance of the model’s output, and reduce the need for further revisions or corrections.
However, sometimes there are scenarios where clarity is not enough. We might need to infer the way of thinking of our LLM to make it more robust with respect to its task. In next section, we are going to examine one of this techniques, very useful in case of complex tasks to accomplish.
Split complex tasks into subtasks
Prompt engineering is a technique that involves designing effective inputs for large language models (LLMs) to perform various tasks. Sometimes, the tasks are too complex or ambiguous for a single prompt to handle, and it is better to split them into simpler subtasks that can be solved by different prompts.
Here are some examples of splitting complex tasks into subtasks:
-
Text summarization: A complex task that involves generating a concise and accurate summary of a long text. This task can be split into subtasks such as:
-
Extracting the main points or keywords from the text.
-
Rewriting the main points or keywords in a coherent and fluent way.
-
Trimming the summary to fit a desired length or format.
-
-
Machine translation: A complex task that involves translating a text from one language to another. This task can be split into subtasks such as:
-
Detecting the source language of the text.
-
Converting the text into an intermediate representation that preserves the meaning and structure of the original text.
-
Generating the text in the target language from the intermediate representation.
-
-
Poem generation: A creative task that involves producing a poem that follows a certain style, theme, or mood. This task can be split into subtasks such as:
-
Choosing a poetic form (such as sonnet, haiku, limerick, etc.) and a rhyme scheme (such as ABAB, AABB, ABCB, etc.) for the poem.
-
Generating a title and a topic for the poem based on the user’s input or preference.
-
Generating the lines or verses of the poem that match the chosen form, rhyme scheme, and topic.
-
Refining and polishing the poem to ensure coherence, fluency, and originality.
-
-
Code generation: A technical task that involves producing a code snippet that performs a specific function or task. This task can be split into subtasks such as:
-
Choosing a programming language (such as Python, Java, C++, etc.) and a framework or library (such as TensorFlow, PyTorch, React, etc.) for the code.
-
Generating a function name and a list of parameters and return values for the code based on the user’s input or specification.
-
Generating the body of the function that implements the logic and functionality of the code.
-
Adding comments and documentation to explain the code and its usage.
-
-
Let’s consider the following example. We will provide the model with a short article and ask it to summarize it following these instructions:
You are an AI assistant that summarize articles.
To complete this task, do the following subtasks:
Read the provided article context comprehensively and identified the main topic and key points
Generated a paragraph summary of the current article context that captures the essential information and conveys the main idea
Print each step of the process.
This is the short article I will provide:
Large Language Models (LLMs), a subset of artificial intelligence, have revolutionized the field of natural language processing by demonstrating an unprecedented ability to understand and generate human-like text. These models are trained on vast datasets comprising diverse linguistic inputs, enabling them to produce coherent and contextually relevant responses across a wide range of topics. By leveraging architectures such as transformers, LLMs like GPT-3 and its successors can complete text, answer questions, perform translations, and even engage in complex dialogue. Their applications span from automated customer support and content creation to advanced research and education tools. Despite their incredible capabilities, LLMs also pose challenges, including the potential for biases inherent in training data and the risk of generating misleading or false information. As the development of LLMs continues to advance, ongoing efforts in ethical AI research and deployment strategies are crucial to harness their benefits responsibly and effectively.
Let’s see how the model works:

Figure 6: Example of OpenAI GPT-4o splitting a task into subtasks to generate a summary.
Splitting complex tasks into easier sub-tasks is a powerful technique, nevertheless it does not address one of the main risks of LLM-generated content, that is, having a wrong output. In next two sections, we are going to see some techniques that are mainly aimed at addressing this risk.
Ask for justification
LLMs are built in such a way that they predict the next token based on the previous ones without looking back at their generations. This might lead the model to output wrong content to the user, yet in a very convincing way. If the LLM-powered application does not provide specific reference to that response, it might be hard to validate the ground truth behind it.
Henceforth, specifying in the prompt to support the LLM’s answer with some reflections and justification could prompt the model to recover from its actions.
Furthermore, asking for justification might be useful also in case of answers that are right, but simply we don’t know the LLM thought process behind it. For example, let’s say we want our LLM to solve riddles. To do so, we can instruct it as follows:
You are an AI assistant specialized in solving riddles.
Given a riddle, solve it the best you can.
Provide a clear justification of your answer and the reasoning behind it.
As you can see, I’ve specified in the metaprompt to the LLM to justify its answer and also socializing its reasoning. Let’s test it with the following riddle:
What has a face and two hands, but no arms or legs?
Output:

Figure 7: Example of OpenAI’s GPT-4o providing justification after solving a riddle.
Justifications are a great tool to make your model more reliable and robust, since they “force” it to re-think about its output, as well as providing us with a view of how the reasoning was set to solve the problem.
With a similar approach, we could also intervene at different prompts level to improve our LLM’s performance. For example, we might discover that the model is systematically tackling a mathematical problem in the wrong way, henceforth we might want to suggest the right approach directly at the metaprompt level. Another example might be that of asking the model to generate multiple outputs – along with their justifications – to evaluate different reasoning techniques and prompt the best one in the metaprompt.
In next section, we are going to focus on one of this example, more specifically the possibility of generating multiple outputs and then picking the most likely one.
Generate many outputs, then use the model to pick the best one
As we saw in the previous section, LLMs are built in such a way that they predict the next token based on the previous ones without looking back at their generations. If this is the case, if one sampled token is the wrong one (in other words, if the model is unlucky), the LLM will keep generating wrong tokens and, henceforth, wrong content. Now, the bad news is that, unlike humans, LLMs cannot recover from errors on their own. It means that, if we ask them, they aknowledge the error, but we need to explicitly prompt them to think about that.
One way to overcome this limitation is that of broadening the space of probabilities of picking the right token. Rather than generating just one response, we can prompt the model to generate multiple response, and then picking the one which is most suitable for the user’s query. This splits the job into two sub-jobs for our LLM:
-
Generating multiple responses to user’s query;
-
Comparing those responses and picking the best one, according to some criteria we can specify in the metaprompt.
Let’s see an example, following up with the riddles examined in the previous section.
You are an AI assistant specialized in solving riddles.
Given a riddle, you have to generate three answers to the riddle.
For each answer, be specific about the reasoning you made.
Then, among the three answer, select the one which is most plausible given the riddle.
In this case, I’ve prompted the model to generate three answers to the riddle, then to give me the most likely, justifying why. Let’s see the result:

Figure 8: Example of GPT-4o generating 3 plausible answers and picking the most likely one, providing justification.
As aforementioned, forcing the model to tackle a problem with different approaches is a way to collect multiple samples of reasonings, which might serve as further instructions in the metaprompt. For example, if we want the model to always propose us something which is not the most straightforward solution to a problem – in other words, if we want it to “think differently” – we might force it to solve a problem in N ways and then use the most creative reasoning as framework in the metaprompt.
The last element we are going to examine is the overall structure we want to give to our metaprompt. In fact, in previous examples, we saw a sample system message with some statements and instructions. In next section, we will see how the order and “strength” of those statements and instructions is not an invariant.
Use delimiters
The last principle to be covered is related to the format we want to give to our meta prompt. This helps our LLM to better understand its intents as well as making relations among sections and paragraphs.
To achieve this, we can use delimiters within our prompt. A delimiter can by any sequence of characters or symbols that is clearly mapping a schema rather than a concept. For example, we can consider delimiters the following sequences:
-
-
====
-
-
`
And so on. Let’s consider, for example, a meta prompt that aims at instructing the model to translate user’s tasks into Python code, providing also an example to do so.
You are a Python expert that produces python code as per user's request.
===>START EXAMPLE
---User Query---
Give me a function to print a string of text.
---User Output---
Below you can find the described function:
```def my_print(text):
#returning the printed text
return print(text)
<===END EXAMPLE
Let’s see how it works:

Figure 9: Sample output of a model using delimiters in the system message.
As you can see, it also printed the code in backsticks as showed within the system message.
All the principles examined up to this point are general rules that can make your LLM-powered application more robust. In next sections, we are going to see some advanced techniques for prompt engineering that address the way the model reason and think to the answer, before providing it to the final user.
## Advanced techniques
In previous sections, we covered some basics techniques of prompt engineering. Those techniques should be kept in mind regardless of the type of application your are developing, since are general best practices that improve your LLM performance anyway.
On the other hand, there are some advanced techniques which might be implemented for specific scenarios, that we are going to cover in the upcoming sections.
### Chain of Thoughts
Introduced in the paper “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” by Wei et al., Chain of thought (CoT) is a technique that enables complex reasoning capabilities through intermediate reasoning steps. It also encourages the model to explain its reasoning, “forcing” it not to be too fact and risking to give the wrong response (as we saw in previous sections).
Let’s say that we want to prompt our LLM to solve first-degree equations. To do so, we are going to provide it with a generic reasoning list as a metaprompt:
To solve a generic first-degree equation, follow these steps:
1. Identify the Equation: Start by identifying the equation you want to solve. It should be in the form of "ax + b = c," where 'a' is the coefficient of the variable, 'x' is the variable, 'b' is a constant, and 'c' is another constant.
2. Isolate the Variable: Your goal is to isolate the variable 'x' on one side of the equation. To do this, perform the following steps:
a. Add or Subtract Constants: Add or subtract 'b' from both sides of the equation to move constants to one side.
b. Divide by the Coefficient: Divide both sides by 'a' to isolate 'x'. If 'a' is zero, the equation may not have a unique solution.
3. Simplify: Simplify both sides of the equation as much as possible.
4. Solve for 'x': Once 'x' is isolated on one side, you have the solution. It will be in the form of 'x = value.'
5. Check Your Solution: Plug the found value of 'x' back into the original equation to ensure it satisfies the equation. If it does, you've found the correct solution.
6. Express the Solution: Write down the solution in a clear and concise form.
7. Consider Special Cases: Be aware of special cases where there may be no solution or infinitely many solutions, especially if 'a' equals zero.
Equation:
Let’s see how it works:

Figure 10: Output of the model solving an equation with CoT approach.
Note that you can also combine combine it with few-shot prompting to get better results on more complex tasks that require reasoning before responding.
With CoT, we are prompting the model to generate intermediate reasoning steps. This is also a component of another reasoning technique that we are going to examine in next section.
### ReAct
Introduced in the paper “ReAct: Synergizing Reasoning and Acting in Language Models” by Yao et al., ReAct (Reason and Act) is a general paradigm that combines reasoning and acting with large language models. ReAct prompts the language model to generate verbal reasoning traces and actions for a task, and also receives observations from external sources such as web search or databases. This allows the language model to perform dynamic reasoning, and quickly adapt its acting plan based on external information. For example, you can prompt the language model to answer a question by first reasoning about the question, then performing an action to send a query to the web, then receiving an observation from the search results, and then continuing with this thought, action, observation loop until it reaches a conclusion.
The difference between chain of thought and ReAct approaches is that chain of thought prompts the language model to generate intermediate reasoning steps for a task, while ReAct prompts the language model to generate intermediate reasoning steps, actions, and observations for a task.
Note that the “action” phase is generally related to the possibility for our LLM to interact with external tools, such as the web search. However, in the following example we won’t use tools, but rather refer to “action” any task we ask the model to do for us.
This is how the ReAct metaprompt might look like:
Answer the following questions as best you can.
Use the following format:
Question: the input question you must answer
Thought: you should always think about what to do
Action: the action to take
Action Input: the input to the action
Observation: the result of the action
... (this Thought/Action/Action Input/Observation can repeat N times)
Thought: I now know the final answer
Final Answer: the final answer to the original input question
Let’s see how does it work with a simple user’s query:

Figure 11: Example of ReAct metaprompting
As you can see, in this scenario the model skipped the Action Input and the following observation, since we didn’t provide it with any tool to be used.
This is a great example of how prompting a model to think step by step and explicit each step of the reasoning makes it “wiser” and cautious before answering. It is also a great technique to prevent hallucination.
Overall, prompt engineering is a powerful discipline, still in its emerging phase yet already widely adopted within LLM-powered applications. In next Chapters, we are going to see concrete applications of this techniques.
## Avoiding the risk of hidden bias and taking into account ethical considerations in ChatGPT
ChatGPT has been provided with the Moderator API so that it cannot engage in conversations that might be unsafe. The Moderator API is a classification model performed by a GPT model based on the following classes: violence, self-harm, hate, harassment, and sex. For this, OpenAI uses anonymized data and synthetic data (in zero-shot form) to create synthetic data.
The Moderation API is based on a more sophisticated version of the content filter model available among OpenAI APIs. We discussed this model in *Chapter 1*, where we saw how it is very conservative toward false positives rather than false negatives.
However, there is something we can refer to as **hidden bias**, which derives directly from the knowledge base the model has been trained on. For example, concerning the main chunk of training data of GPT-3, known as the **Common Crawl**, experts believe that it was written mainly by white males from Western countries. If this is the case, we are already facing a hidden bias of the model, which will inevitably mimic a limited and unrepresentative category of human beings.
In their paper, *Languages Models are Few-Shots Learners*, OpenAI’s researchers Tom Brown et al. ([https://arxiv.org/pdf/2005.1416](https://arxiv.org/pdf/2005.1416)) created an experimental setup to investigate racial bias in GPT-3\. The model was prompted with phrases containing racial categories and 800 samples were generated for each category. The sentiment of the generated text was measured using Senti WordNet based on word co-occurrences on a scale ranging from -100 to 100 (with positive scores indicating positive words, and vice versa).
The results showed that the sentiment associated with each racial category varied across different models, with Asian consistently having a high sentiment and Black consistently having a low sentiment. The authors caution that the results reflect the experimental setup and that socio-historical factors may influence the sentiment associated with different demographics. The study highlights the need for a more sophisticated analysis of the relationship between sentiment, entities, and input data:

Figure 4.14 – Racial sentiment across models
This hidden bias could generate harmful responses not in line with responsible AI principles.
However, it is worth noticing how ChatGPT, as well as all OpenAI models, are subject to continuous improvements. This is also consistent with OpenAI’s AI **alignment** ([https://openai.com/ alignment/](https://openai.com/)), whose research focuses on training AI systems to be helpful, truthful, and safe.
For example, if we ask ChatGPT to formulate guesses based on people’s gender and race, it will not accommodate the exact request, but rather provide us with hypothetical function as well as a huge disclaimer:

Figure 4.15 – Example of GPT-4o improving over time since it gives an unbiased response
Overall, despite the continuous improvement in the domain of ethical principles, while using ChatGPT, we should always make sure that the output is in line with those principles and not biased.
The concepts of bias and ethics within ChatGPT and OpenAI models have a wider collocation within the whole topic of responsible AI, which we are going to focus on in the last chapter of this book.
## Summary
In this chapter, we have dived deeper into the concept of prompt design and engineering since it’s the most powerful way to control the output of ChatGPT and LLMs in general. We learned how to leverage different levels of shot learning to make LLMs more tailored toward our objectives.
We started with an introduction to the concept of prompt engineering and why it is important, to then move towards the basic principles – including clear instructions, asking for justification etc.
Then, we moved towards more advanced techniques, which are meant to shape the reasoning approach of our LLM: few-shot learning, CoT and ReAct.
Prompt engineering is an emerging discipline which is paving the way for a new category of applications, infused with LLMs. In next chapters, we will see those techniques in action building real-world applications using LLMs.
Starting from the next chapter, we will dive deeper into different domains where ChatGPT can boost productivity and have a disruptive impact on the way we work today.
## References
The following are the references for this chapter:
* [https://arxiv.org/abs/2005.14165](https://arxiv.org/abs/2005.14165)
* [https://dl.acm.org/doi/10.1145/3442188.3445922](https://dl.acm.org/doi/10.1145/3442188.3445922)
* [https://openai.com/alignment/](https://openai.com/alignment/)
* [https://twitter.com/spiantado/status/1599462375887114240?ref](https://twitter.com/spiantado/status/1599462375887114240?ref)
* ReAct approach. [https://arxiv.org/abs/2210.03629](https://arxiv.org/abs/2210.03629)
* Chain of Thoughts approach. [https://arxiv.org/abs/2201.11903](https://arxiv.org/abs/2201.11903)
* What is prompt engineering. [https://www.mckinsey.com/featured-insights/mckinsey-explainers/what-is-prompt-engineering](https://www.mckinsey.com/featured-insights/mckinsey-explainers/what-is-prompt-engineering)
* Prompt engineering techniques. [https://blog.mrsharm.com/prompt-engineering-guide/](https://blog.mrsharm.com/prompt-engineering-guide/)
* Prompt engineering principles. [https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/advanced-prompt-engineering?pivots=programming-language-chat-completions](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/advanced-prompt-engineering?pivots=programming-language-chat-completions)
# 4 Boosting Day-to-Day Productivity with ChatGPT
## Join our book community on Discord

[https://packt.link/EarlyAccess](https://packt.link/EarlyAccess)
In this chapter, we will cover the main activities ChatGPT can perform for general users daily to boost their productivity. The chapter will focus on concrete examples of writing assistance, decision- making, information retrieval, and so on, with suggestions and prompts so that you can implement them on your own.
By the end of this chapter, you will have learned how to use ChatGPT as a booster for the following activities:
* Daily activities such as organizing agendas, meal-prepping, grocery shopping, and so on
* Generating brand-new text content
* Improving your writing skills and adapting the same content to different audiences
* Retrieving documentation and information for research and competitive intelligence
## Technical requirements
For this chapter, you will require a ChatGPT account.
## ChatGPT as a daily assistant
ChatGPT can serve as a valuable daily assistant, helping you manage your tasks and streamline your workflow. It can optimize your daily routine by providing personalized assistance, thus saving you time and enhancing your efficiency.
Let’s start with a general suggestion on how I could make my day more productive:

Figure 5.1 – An example of ChatGPT generating a productive routine
Another interesting usage of ChatGPT in organizing my week is that I can use it as a meal prep assistant. Note that here I also asked for a specific table formatting:

Figure 5.3 – Meal prep for my working week generated by ChatGPT
Alongside meal prepping, ChatGPT can also generate a workout plan according to my requirements:

Figure 5.4 – Workout plan generated by ChatGPT
ChatGPT can also be a loyal and disciplined study partner. It could help you, for example, in summarizing long papers so that you gain a first overview of the discussed topic or help you prepare for exams.
Namely, let’s say you are preparing for mathematics exams using the university book titled *Mathematics for Economics and Business*, by Lorenzo Peccati et al. Before diving deeper into each chapter, you might want to get an overview of the content and main topics that are discussed, as well as whether you need further prerequisites and – mostly importantly if I’m preparing for an exam – how long it will take to study it. You can ask ChatGPT for this:

Figure 5.5 – ChatGPT providing an overview of a university book
When it comes to questions about specific assets or personal information, we are increasing the risk of ChatGPT’s hallucination. In fact, in the previous example, the model was able to answer since the synopsis of the book was clearly part of the training set, yet the model didn’t have full visibility on the book, as it’s not available on the web for free. If this is the case, it might be a good practice to specify the model to navigate the web to answer these specific questions. Let’s see an example of it:

Figure 1: Example of ChatGPT invoking the web plug-in.
As you can see, now ChatGPT invoked the web plug-in and searched over 5 files, also providing the links it navigated through.
You can also ask ChatGPT to ask you some questions about the material you have just studied:

Figure 5.6 – Example of ChatGPT acting as a professor
> The technique of *Act as*… is a great example of efficient prompting techniques, and it can be listed among the examples described in *Chapter 3*.
Now, let’s look at some more examples of using ChatGPT for more specific tasks, including text generation, writing assistance, and information retrieval.
## Generating text
As a language model, ChatGPT is particularly suited for generating text based on users’ instructions. For example, you could ask ChatGPT to generate emails, drafts, or templates that target a specific audience:

Figure 5.7 – Example of an email generated by ChatGPT
Another example might be asking ChatGPT to create a pitch structure for a presentation you have to prepare (I’ll put here only the first slide as an example):

Figure 5.8 – Slideshow agenda and structure generated by ChatGPT
You can also generate blog posts or articles about trending topics this way. Here is an example:

Figure 5.9 – Example of a blog post with relevant tags and SEO keywords generated by ChatGPT
We can even get ChatGPT to reduce the size of the post to make it fit for a tweet. Here is how we can do this:

Figure 5.10 – ChatGPT shrinks an article into a Twitter post
Finally, ChatGPT can also generate video or theatre scripts, including the scenography and the suggested editing. The following figure shows an example of a theatre dialog including scenography and actors:

Figure 5.11 – Theatre dialog with scenography generated by ChatGPT
I only provided one scene out of the four generated by ChatGPT, to keep you in suspense regarding the ending…
Overall, whenever new content needs to be generated from scratch, ChatGPT does a very nice job of providing a first draft, which could act as the starting point for further refinements.
However, ChatGPT can also support pre-existing content by providing writing assistance and translation, as we will see in the next section.
## Improving writing skills and translation
Sometimes, rather than generating new content, you might want to revisit an existing piece of text. It this be for style improvement purposes, audience changes, language translation, and so on.
Let’s look at some examples. Imagine that I drafted an email to invite a customer of mine to a webinar. I wrote two short sentences. Here, I want ChatGPT to improve the form and style of this email since the target audience will be executive level:

Figure 5.12 – Example of an email revisited by ChatGPT to target an executive audience
Now, let’s ask the same thing but with a different target audience:

Figure 5.13 – Example of the same email with a different audience, generated by ChatGPT
ChatGPT can also give you some feedback about your writing style and structure.
Imagine, for example, that you wrote an introduction for an essay titled *The History of Natural Language Processing* and you want some feedback about the writing style and its consistency with the title:

Figure 5.15 – Example of ChatGPT giving feedback on an introduction for an essay
As you can see, not only ChatGPT provided me with feedback and tips for the overall essay, but also it gave me a revised introduction which incorporates all of them.
Let’s also ask ChatGPT to be more specific regarding its suggestion of “Refining Examples”:

Figure 5.16 – Example of ChatGPT elaborating on something it mentioned
I’m also interested in knowing whether my introduction was consistent with the title or whether I’m taking the wrong direction:

Figure 5.17 – ChatGPT provides feedback about the consistency of the introduction with the title
I was impressed by this last one. ChatGPT was smart enough to see that there was no specific mention of the history of NLP in my introduction. Nevertheless, it sets up the expectation about that topic to be treated later on. This means that ChatGPT also has expertise in terms of how an essay should be structured and it was very precise in applying its judgment, knowing that it was just an introduction.
Let’s now unveil the last ChatGPT skill of the Chapter. In fact, ChatGPT is also an excellent tool for translation. It knows at least 95 languages (if you have doubts about whether your language is supported, you can always ask ChatGPT directly). Here, however, there is a consideration that might arise: what is the added value of ChatGPT for translation when we already have cutting-edge tools such as Google Translate?
To answer this question, we have to consider some key differentiators and how we can leverage ChatGPT’s embedded translations capabilities:
* ChatGPT can capture the intent. This means that you could also bypass the translation phase since it is something that ChatGPT can do in the backend. For example, if you write a prompt to produce a social media post in French, you could write that prompt in any language you want – ChatGPT will automatically detect it (without the need to specify it in advance) and understand your intent:

Figure 5.18 – Example of ChatGPT generating an output in a language that is different from the input
* ChatGPT can capture the more refined meaning of slang or idioms. This allows for a translation that is not literal so that it can preserve the underlying meaning. Namely, let’s consider the British expression It’s not my cup of tea, which indicates that something is not to one’s liking or preference. Let’s ask both ChatGPT and Google Translate to translate it into Italian:

Figure 5.19 – Comparison between ChatGPT and Google Translate while translating from English into Italian
As you can see, ChatGPT can provide several Italian idioms that are equivalent to the original one, also in their slang format. On the other hand, Google Translate performed a literal translation, leaving behind the real meaning of the idiom.
* As with any other task, you can always provide context to ChatGPT. So, if you want your translation to have a specific slang or style, you can always specify it in the prompt. Or, even funnier, you can ask ChatGPT to translate your prompt with a sarcastic touch:

Figure 5.20 – Example of ChatGPT translating a prompt with a sarcastic touch. The original content of the prompt was taken from OpenAI’s Wikipedia page: https://it.wikipedia.org/wiki/OpenAI
All these scenarios highlight one of the key killing features of ChatGPT and OpenAI models in general. Since they represent the manifestation of what OpenAI defined as **Artificial General Intelligence** (**AGI**), they are not meant to be specialized (that is, constrained) on a single task. On the contrary, they are meant to serve multiple scenarios dynamically so that you can address a wide range of use cases with a single model.
In conclusion, ChatGPT is able not only to generate new text but also to manipulate existing material to tailor it to your needs. It has also proven to be very precise at translating between languages, also keeping the jargon and language-specific expressions intact.
In the next section, we will see how ChatGPT can assist us in retrieving information and competitive intelligence.
## Quick information retrieval and competitive intelligence
Information retrieval and competitive intelligence are yet other fields where ChatGPT is a game-changer. The very first example of how ChatGPT can retrieve information is the most popular way it is used right now: as a search engine. Every time we ask ChatGPT something, it can retrieve information from its knowledge base and reframe it in an original way.
One example involves asking ChatGPT to provide a quick summary or review of a book we might be interested in reading:

Figure 5.21 – Example of ChatGPT providing a summary and review of a book
Alternatively, we could ask for some suggestions for a new book we wish to read based on our preferences:

Figure 5.22 – Example of ChatGPT recommending a list of books, given my preferences
Furthermore, if we design the prompt with more specific information, ChatGPT can serve as a tool for pointing us toward the right references for our research or studies.
Namely, you might want to quickly retrieve some background references about a topic you want to learn more about – for example, feedforward neural networks. Something you might ask ChatGPT is to point you to some websites or papers where this topic is widely treated:

Figure 5.23 – Example of ChatGPT listing relevant references
As you can see, ChatGPT was able to provide me with relevant references to start studying the topic. However, it could go even further in terms of competitive intelligence.
Let’s consider I’m writing a book titled Introduction to *Convolutional Neural Networks – an Implementation with Python*. I want to do some research about the potential competitors in the market. The first thing I want to investigate is whether there are already some competitive titles around, so I can ask ChatGPT to generate a list of existing books with the same content:

Figure 5.24 – Example of ChatGPT providing a list of competitive books
You can also ask for feedback in terms of the saturation of the market you want to publish in:

Figure 5.25 – ChatGPT advising about how to be competitive in the market
Finally, let’s ask ChatGPT to be more precise about what I should do to be competitive in the market where I will operate:

Figure 5.26 – Example of how ChatGPT can suggest improvements regarding your book content to make it stand out
ChatGPT was pretty good at listing some good tips to make my book unique.
Overall, ChatGPT can be a valuable assistant for information retrieval and competitive intelligence. However, it is important to remember the knowledge base cut-off is 2021: this means that, whenever we need to retrieve real-time information, or while making a competitive market analysis for today, we might not be able to rely on ChatGPT.
Nevertheless, this tool still provides excellent suggestions and best practices that can be applied, regardless of the knowledge base cut-off.
## Summary
All the examples we saw in this chapter were modest representations of what you can achieve with ChatGPT to boost your productivity. These small hacks can greatly assist you with activities that might be repetitive (answering emails with a similar template rather than writing a daily routine) or onerous (such as searching for background documentation or competitive intelligence).
In the next chapter, we are going to dive deeper into three main domains where ChatGPT is changing the game – development, marketing, and research.
# 5 Developing the Future with ChatGPT
## Join our book community on Discord

[https://packt.link/EarlyAccess](https://packt.link/EarlyAccess)
In this chapter, we will discuss how developers can leverage ChatGPT. The chapter focuses on the main use cases ChatGPT addresses in the domain of developers, including code review and optimization, documentation generation, and code generation. The chapter will provide examples and enable you to try the prompts on your own.
After a general introduction about the reasons why developers should leverage ChatGPT as a daily assistant, we will focus on ChatGPT and how it can do the following:
* Why ChatGPT for developers?
* Generate, optimize, and debug code
* Generate code-related documentation and debug your code
* Explain **machine learning** (**ML**) models to help data scientists and business users with model interpretability
* Translate different programming languages
By the end of this chapter, you will be able to leverage ChatGPT for coding activities and use it as an assistant for your coding productivity.
## Why ChatGPT for developers?
Personally, I believe that one of the most mind-blowing capabilities of ChatGPT is that of dealing with code. Of any type. We’ve already seen in previous chapters some examples of ChatGPT generating Python code. However, ChatGPT capabilities for developers go way beyond that example. It can be a daily assistant for code generation, explanation, and debugging.
Among the most popular languages, we can certainly mention Python, JavaScript, SQL, and C#. However, ChatGPT covers a wide range of languages, as disclosed by itself:

Figure 6.1 – ChatGPT lists the programming languages it is able to understand and generate
Whether you are a backend/frontend developer, a data scientist, or a data engineer, whenever you work with a programming language, ChatGPT can be a game changer, and we will see how in the several examples in the next sections.
From the next section onward, we will dive deeper into concrete examples of what ChatGPT can achieve when working with code. We will see end-to-end use cases covering different domains so that we will get familiar with using ChatGPT as a code assistant.
## Generating, optimizing, and debugging code
The primary capability you should leverage is ChatGPT code generation. How many times have you been looking for a pre-built piece of code to start from? Generating the utils functions, sample datasets, SQL schemas, and so on? ChatGPT is able to generate code based on input in natural language:

Figure 6.2 – Example of ChatGPT generating a Python function to write into CSV files
As you can see, not only was ChatGPT able to generate the function, but also it was able to explain what the function does, how to use it, and what to substitute with generic placeholders such as `my_folder`.
Now let’s raise the difficulty bar. If ChatGPT is capable of generating a Python function, could it generate an entire Video Game as well? Let’s try. What I want to do, is providing ChatGPT with an illustration of the type of game I want to develop, and ask it to replicate it with code. The following is an illustration of my desired game (can you guess the name?):

Figure 1: Illustration of the game Pac-Man.
Now let’s ask ChatGPT to reproduce it:

Figure 2: Example of ChatGPT generating a HTML, CSS and JS code.
As disclaimed by ChatGPT, the full game required a lot of code, however, let’s see how the generated code works insofar (to run the code, I used the online tool codepen.io):

Figure 3: Pac-Man game generated by ChatGPT.
As you can see, the draft product looks already similar to what I’m aiming for! This is an example on how Generative AI can help you overcoming the “curse” of starting from scratch: in fact, starting from a whitepaper can be sometimes blocking, while having a draft product to start from can not only speed up the overall process, but also stimulate creativity and improve the quality of the result.
ChatGPT can also be a great assistant for code optimization. In fact, it might save us some running time or compute power to make optimized scripts starting from our input. This capability might be compared, in the domain of natural language, to the writing assistance feature we saw in Chapter 5 in the Improving writing skills and translation section.
For example, imagine you want to create a list of odd numbers starting from another list. To achieve the result, you wrote the following Python script (for the purpose of this exercise, I will also track the execution time with the timeit and datetime libraries):
from timeit import default_timer as timer from datetime import timedelta
start = timer()
elements = list(range(1_000_000)) data = []
for el in elements: if not el % 2:
if odd number data.append(el)
end = timer() print(timedelta(seconds=end-start))
Let’s see how long it takes to run:

Figure 4: Speed of execution of a Python function.
The execution time was 00.115022 seconds. What happens if we ask ChatGPT to optimize this script?

Figure 6.6 – ChatGPT generating optimized alternatives to a Python script
ChatGPT provided me with two examples to achieve the same results with lower execution time.
Let’s test both of them in a Jupyter Notebook:
2023-03-25 11:27:10.270 Uncaught app exception Traceback (most recent call last):
File "C:\Users\vaalt\Anaconda3\lib\site-packages\streamlit\runtime\ scriptrunner\script_runner.py", line 565, in _run_script
exec(code, module. dict )
File "C:\Users\vaalt\OneDrive\Desktop\medium articles\llm.py", line 129, in
user_input = get_text()
File "C:\Users\vaalt\OneDrive\Desktop\medium articles\llm.py", line 50, in get_text
input_text = st.text_input("You: ", st.session_state['input'], key='input', placeholder = 'Your AI assistant here! Ask me anything...', label_visibility = 'hidden')
File "C:\Users\vaalt\Anaconda3\lib\site-packages\streamlit\runtime\ metrics_util.py", line 311, in wrapped_func
result = non_optional_func(*args, **kwargs)
File "C:\Users\vaalt\Anaconda3\lib\site-packages\streamlit\elements\ text_widgets.py", line 174, in text_input
return self._text_input(

Figure 6.7 – Speed of execution of two alternative functions generated by ChatGPT.
As you can see, both methods lead to a great reduction in time, respectively of 44.30% and 20.68%.
On top of code generation and optimization, ChatGPT can also be leveraged for error explanation and debugging. Sometimes, errors are difficult to interpret; hence a natural language explanation can be useful to identify the problem and drive you toward the solution.
For example, while running a .py file from my command line, I get the following error:
File "C:\Users\vaalt\Anaconda3\lib\site-packages\streamlit\elements\ text_widgets.py", line 266, in _text_input
text_input_proto.value = widget_state.value
TypeError: [] has type list, but expected one of: bytes, Unicode
Let’s see whether ChatGPT is able to let me understand the nature of the error. To do so, I simply provide ChatGPT with the text of the error and ask it to give me an explanation.

Figure 6.8 – ChatGPT explaining a Python error in natural language
Finally, let’s imagine I wrote a function in Python that takes a string as input and returns the same string with an underscore after each letter.
In the preceding example, I was expecting to see the g_p_t_ result; however, it only returned t_
with this code:

Figure 6.9 – Bugged Python function
Let’s ask ChatGPT to debug this function for us:

Figure 6.10 – Example of ChatGPT debugging a Python function
Impressive, isn’t it? Again, ChatGPT provided the correct version of the code, and it helped in the explanation of where the bugs were and why they led to an incorrect result. Let’s see whether it works now:

Figure 6.11 – Python function after ChatGPT debugging
Well, it obviously does!
These and many other code-related functionalities could really boost your productivity, shortening the time to perform many tasks.
However, ChatGPT goes beyond pure debugging. Thanks to the incredible language understanding of the GPT model behind, this GenAI tool is able to generate proper documentation alongside the code, as well as explain exactly what a string of code will do, which we will see in the next section.
## Generating documentation and code explainability
Whenever working with new applications or projects, it is always good practice to correlate your code with documentation. It might be in the form of a docstring that you can embed in your functions or classes so that others can invoke them directly in the development environment.
For example, let’s consider the same function developed in the previous section and let’s make it a Python class:
class UnderscoreAdder:
def init(self, word):
self.word = word
def add_underscores(self):
new_word = ""
for i in range(len(self.word)):
new_word += self.word[i] + "_"
return new_word
We can test it as follows:

Figure 5: Testing the UnderscoreAdder class.
Now, let’s say I want to be able to retrieve the docstring documentation using the `UnderscoreAdder` ? convention. By doing so with Python packages, functions, and methods, we have full documentation of the capabilities of that specific object, as follows (an example with the pandas Python library):
Figure 6.13 – Example of the pandas library documentation
So, let’s now ask ChatGPT to produce the same result for our `UnderscoreAdder` class.

Figure 6.14 – ChatGPT updating the code with documentation
As a result, if we update our class as shown in the preceding code and `UnderscoreAdder`?, we will get the following output:

Figure 6.15 – The new UnderscoreAdder class documentation
Finally, ChatGPT can also be leveraged to explain what a script, function, class, or other similar things do in natural language. We have already seen many examples of ChatGPT enriching its code-related response with clear explanations. However, we can boost this capability by asking specific questions in terms of code understanding.
For example, let’s ask ChatGPT to explain to us what the following Python script does:

Figure 6.16 – Example of ChatGPT explaining a Python script
Code explainability can also be part of the preceding mentioned documentation, or it can be used among developers who might want to better understand complex code from other teams or (as sometimes happens to me) remember what they wrote some time ago.
Thanks to ChatGPT and the capabilities mentioned in this section, developers can easily keep track of the project life cycle in natural language so that it is easier for both new team members and non-technical users to understand the work done so far.
We will see in the next section how code explainability is a pivotal step for ML model interpretability in data science projects.
## Understanding ML model interpretability
Model interpretability refers to the degree of ease with which a human can comprehend the logic behind the ML model’s predictions. Essentially, it is the capability to comprehend how a model arrives at its decisions and which variables are contributing to its forecasts.
Let’s see an example of model interpretability using a deep learning **convolutional neural network** (**CNN**) for image classification. I built my model in Python using Keras. For this purpose, I will download the CIFAR-10 dataset directly from keras.datasets: it consists of 60,000 32x32 color images (so 3-channels images) in 10 classes (airplane, automobile, bird, cat, deer, dog, frog, horse, ship, and truck), with 6,000 images per class. Here, I will share just the body of the model; you can find all the related code in the book’s GitHub repository for data preparation and pre-processing at [https://github.com/PacktPublishing/The-Ultimate-Guide-to-ChatGPT-and-OpenAI/tree/main/Chapter%206%20-%20ChatGPT%20for%20Developers/code](https://github.com/PacktPublishing/The-Ultimate-Guide-to-ChatGPT-and-OpenAI/tree/main/Chapter%206%20-%20ChatGPT%20for%20Developers/code):
model=tf.keras.Sequential()
model.add(tf.keras.layers.Conv2D(32,kernel_ size=(3,3),activation='relu',input_shape=
(32,32,1)))
model.add(tf.keras.layers.MaxPooling2D(pool_size=(2,2))) model.add(tf.keras.layers.Flatten()) model.add(tf.keras.layers.Dense(1024,activation='relu')) model.add(tf.keras.layers.Dense(10,activation='softmax'))
The preceding code is made of several layers that perform different actions. I might be interested in having an explanation of the structure of the model as well as the purpose of each layer. Let’s ask ChatGPT for some help with that (below you can see an extract of the response):

Figure 6.17 – Model interpretability with ChatGPT
As you can see in the preceding figure, ChatGPT was able to give us a clear explanation of the structure and layers of our CNN. It also adds some comments and tips, such as the fact that using the max pooling layer helps reduce the dimensionality of the input.
I can also be supported by ChatGPT in interpreting model results in the validation phase. So, after splitting data into training and test sets and training the model on the training set, I want to see its performance on the test set:

Figure 6.18 – Evaluation metrics
Let’s also ask ChatGPT to elaborate on our validation metrics (truncated output):

Figure 6.19 – Example of ChatGPT explaining evaluation metrics
Once again, the result was really impressive, and it provided clear guidance on how to set up ML experiments in terms of training and test sets. It explains how important it is for the model to be sufficiently generalized so that it does not overfit and is able to predict accurate results on data that it has never seen before.
There are many reasons why model interpretability is important. A pivotal element is that it reduces the gap between business users and the code behind models. This is key to enabling business users to understand how a model behaves, as well as translate it into code business ideas.
Furthermore, model interpretability enables one of the key principles of responsible and ethical AI, which is the transparency of how a model behind AI systems thinks and behaves. Unlocking model interpretability means detecting potential biases or harmful behaviors a model could have while in production and consequently preventing them from happening.
Overall, ChatGPT can provide valuable support in the context of model interpretability, generating insights at the row level, as we saw in the previous example.
The next and last ChatGPT capability we will explore will be yet another boost for developers’ productivity, especially when various programming languages are being used within the same project.
## Translation among different programming languages
In *Chapter 5*, we saw how ChatGPT has great capabilities for translating between different languages. What is really incredible is that natural language is not its only object of translation. In fact, ChatGPT is capable of translating between different programming languages while keeping the same output as well as the same style (namely, it preserves docstring documentation if present).
There are so many scenarios when this could be a game changer.
For example, you might have to learn a new programming language or statistical tool you’ve never seen before because you need to quickly deliver a project on it. With the help of ChatGPT, you can start programming in your language of preference and then ask it to translate to the desired language, which you will be learning alongside the translation process.
Imagine that the project needs to be delivered in MATLAB (a proprietary numerical computing and programming software developed by MathWorks), yet you’ve always programmed in Python. The project consists of classifying images from the **Modified National Institute of Standards and Technology** (**MNIST**) dataset (the original dataset description and related paper can be found here at [http://yann.lecun.com/exdb/mnist/).](http://yann.lecun.com/exdb/mnist/).) The dataset contains numerous handwritten digits and is frequently utilized to teach various image processing systems.
To start, I wrote the following Python code to initialize a deep-learning model for classification:
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras import layers
Load the MNIST dataset
(x_train, y_train), (x_test, y_test) = keras.datasets.mnist.load_ data()
Preprocess the data
x_train = x_train.reshape(-1, 2828) / 255.0 x_test = x_test.reshape(-1, 2828) / 255.0 y_train = keras.utils.to_categorical(y_train) y_test = keras.utils.to_categorical(y_test)
Define the model architecture model = keras.Sequential([
layers.Dense(256, activation='relu', input_shape=(28*28,)), layers.Dense(128, activation='relu'),
layers.Dense(10, activation='softmax')
])
Compile the model
model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
Train the model
history = model.fit(x_train, y_train, validation_split=0.2, epochs=10, batch_size=128)
Evaluate the model
test_loss, test_acc = model.evaluate(x_test, y_test, verbose=0) print('Test accuracy:', test_acc)
Let’s now see what happens if we give the preceding code as context to ChatGPT and ask it to translate it into MATLAB:

Figure 20 – ChatGPT translates Python code into MATLAB
Let’s also see whether it is capable of translating it into other languages such as JavaScript (note that I didn’t have to repeat the code, since it was already in the chat context):

Figure 6.21 – ChatGPT translates Python code into JavaScript
Code translation could also reduce the skill gap between new technologies and current programming capabilities.
Another key implication of code translation is **application modernization**. Indeed, imagine you want to refresh your application stack, namely migrating to the cloud. You could decide to initiate with a simple lift and shift going toward **Infrastructure-as-a-Service** (**IaaS**) instances (such as Windows or Linux **virtual machines** (**VMs**)). However, in a second phase, you might want to refactor, rearchitect, or ever rebuild your applications.
The following diagram depicts the various options for application modernization:

Figure 6.22 – Four ways you can migrate your applications to the public cloud
ChatGPT and OpenAI Codex models can help you with the migration. Consider mainframes, for example.
Mainframes are computers that are predominantly employed by large organizations to carry out essential tasks such as bulk data processing for activities such as censuses, consumer and industry statistics, enterprise resource planning, and large-scale transaction processing. The application programming language of the mainframe environment is **Common Business Oriented Language** (**COBOL**). Despite being invented in 1959, COBOL is still in use today and is one of the oldest programming languages in existence.
As technology continues to improve, applications residing in the realm of mainframes have been subject to a continuous process of migration and modernization aimed at enhancing existing legacy mainframe infrastructure in areas such as interface, code, cost, performance, and maintainability.
Of course, this implies translating COBOL to more modern programming languages such as C# or Java. The problem is that COBOL is unknown to most of the new-generation programmers; hence there is a huge skills gap in this context.
Let’s consider a COBOL script that reads an input number, adds 10 to it, and then prints the result.
IDENTIFICATION DIVISION.
PROGRAM-ID. AddTen.
DATA DIVISION.
WORKING-STORAGE SECTION.
01 INPUT-NUMBER PIC 9(5).
01 RESULT-NUMBER PIC 9(5).
PROCEDURE DIVISION.
DISPLAY 'Enter a number: '.
ACCEPT INPUT-NUMBER.
COMPUTE RESULT-NUMBER = INPUT-NUMBER + 10.
DISPLAY 'Result after adding 10: ' RESULT-NUMBER.
STOP RUN.
I then passed the previous COBOL script to ChatGPT so that it can use it as context to formulate its response. Let’s now ask ChatGPT to translate that script into C#:

Figure 6.23 – Example of ChatGPT translating COBOL to C#.
Tools such as ChatGPT can help in reducing the skill gap in this and similar scenarios by introducing a layer that knows both the past and the future of programming.
In conclusion, ChatGPT can be an effective tool for application modernization, providing code upgrading in addition to valuable insights and recommendations for enhancing legacy systems. With its advanced language processing capabilities and extensive knowledge base, ChatGPT can help organizations streamline their modernization efforts, making the process faster, more efficient, and more effective.
## Summary
ChatGPT can be a valuable resource for developers looking to enhance their skills and streamline their workflows. We started by seeing how ChatGPT can generate, optimize, and debug your code, but we also covered further capabilities such as generating documentation alongside your code, explaining your ML models, and translating between different programming languages for application modernization.
Whether you’re a seasoned developer or just starting out, ChatGPT offers a powerful tool for learning and growth, reducing the gap between code and natural language.
In the next chapter, we will dive deeper into another domain of application where ChatGPT could be a game changer: marketing.
# 6 Mastering Marketing with ChatGPT
## Join our book community on Discord

[https://packt.link/EarlyAccess](https://packt.link/EarlyAccess)
In this chapter, we will focus on how marketers can leverage ChatGPT, looking at the main use cases of ChatGPT in this domain, and how marketers can leverage it as a valuable assistant.
We will learn how ChatGPT can assist in the following activities:
* Marketers’ need for ChatGPT
* New product development and the go-to-market strategy
* A/B testing for marketing comparison
* Making more efficient websites and posts with **Search Engine Optimization** (**SEO**)
* Sentiment analysis of textual data
By the end of this chapter, you will be able to leverage ChatGPT for marketing-related activities and to boost your productivity.
## Technical requirements
You will need an OpenAI account to access ChatGPT.
## Marketers’ need for ChatGPT
Marketing is probably the domain where ChatGPT and OpenAI models’ creative power can be leveraged in their purest form. They can be practical tools to support creative development in terms of new products, marketing campaigns, search engine optimization, and so on. Overall, marketers automate and streamline many aspects of their work, while also improving the quality and effectiveness of their marketing efforts.
Here is an example. One of the most prominent and promising use cases of ChatGPT in marketing is personalized marketing. ChatGPT can be used to analyze customer data and generate personalized marketing messages that resonate with individual customers. For example, a marketing team can use ChatGPT to analyze customer data and develop targeted email campaigns that are tailored to specific customer preferences and behavior. This can increase the likelihood of conversion and lead to greater customer satisfaction. By providing insights into customer sentiment and behavior, generating personalized marketing messages, providing personalized customer support, and generating content, ChatGPT can help marketers deliver exceptional customer experiences and drive business growth.
This is one of many examples of ChatGPT applications in marketing. In the following sections, we will look at concrete examples of end-to-end marketing projects supported by ChatGPT.
## New product development and the go-to-market strategy
The first way you can introduce ChatGPT into your marketing activity might be as an assistant in new product development and **go-to-market** (**GTM**) strategy.
In this section, we will look at a step-by-step guide on how to develop and promote a new product. You already own a running clothing brand called **RunFast** and so far you have only produced shoes, so you want to expand your business with a new product line. We will start by brainstorming ideas to create a GTM strategy. Of course, everything is supported by ChatGPT:
* Brainstorming ideas: The first thing ChatGPT can support you with is brainstorming and drafting options for your new product line. It will also provide the reasoning behind each suggestion. So, let’s ask what kind of new product line I should focus on:

Figure 7.1 – Example of new ideas generated by ChatGPT
Out of the three suggestions, we will pick the second one, because of the positive environmental impact that we can make with it, as well as improving our brand reputation. More specifically, I will start with eco-friendly running socks.
* Product name: Now that we have our idea fixed in mind, we need to think of a catchy name for it. Again, I will ask ChatGPT for more options so that I can then pick my favorite one:

Figure 7.2 – A list of potential product names
GreenStride sounds good enough for me – I’ll go ahead with that one.
* Generating catchy slogans: On top of the product name, I also want to share the intent behind the name and the mission of the product line, so that my target audience is captured by it. I want to inspire trust and loyalty in my customers and for them to see themselves reflected in the mission behind my new product line.

Figure 7.3 – A list of slogans for my new product name
Great – now I’m satisfied with the product name and slogan that I will use later on to create a unique social media announcement. Before doing that, I want to spend more time on market research for the target audience.

Figure 7.4 – List of groups of target people to reach with my new product line
It’s important to have in mind different clusters within your audience so that you can differentiate the messages you want to give. In my case, I want to make sure that my product line will address different groups of people, such as competitive runners, casual runners, and fitness enthusiasts.
* Product variants and sales channels: According to the preceding clusters of potential customers, I could generate product variants so that they are more tailored toward specific audiences:

Figure 7.5 – Example of variants of the product line
Similarly, I can also ask ChatGPT to suggest different sales channels for each of the preceding groups:

Figure 7.6 – Suggestions for different sales channels by ChatGPT
* Standing out from the competition: I want my product line to stand out from the competition and emerge in a very saturated market – I want to make it unique. With this purpose in mind, I asked ChatGPT to include social considerations such as sustainability and inclusivity. Let’s ask ChatGPT for some suggestions in that respect:

Figure 7.7 – Example of outstanding features generated by ChatGPT
As you can see, it was able to generate interesting features that could make my product line unique.
* Product description: Now it’s time to start building our Go-To-Market plan. First of all, I want to generate a product description for my website, including all the earlier unique differentiators.

Figure 7.8 – Example of description and SEO keywords generated by ChatGPT
* Fair price: Another key element is determining a fair price for our product. As I differentiated among product variants for different audiences (competitive runners, casual runners, and fitness enthusiasts), I also want to have a price range that takes into account this clustering. Note that, in the following example, ChatGPT is invoking the web search plug-in to retrieve updated information about the current marketplace of running socks in terms of pricing.

Figure 7.9 – Price ranges for product variants
We are almost there. We have gone through many new product development and go-to-market steps, and in each of them, ChatGPT acted as a great support tool.
As one last thing, we can ask ChatGPT to generate an Instagram post about our new product, including relevant hashtags and SEO keywords. We can then generate the image DALL-E, which comes as an embedded plug-in in ChatGPT Plus.

Figure 7.10 – Social media post generated by ChatGPT
And, with the special contribution of DALL-E:

Figure 1: Example of an illustration generated by ChatGPT powered by DALL-E 3.
Here it is the final result:

Figure 7.11 – Instagram post entirely generated by ChatGPT and DALL-E
Of course, many elements are missing here for complete product development and go-to-market. Yet, with the support of ChatGPT (and the special contribution of DALL-E 3) we managed to brainstorm a new product line and variants, potential customers, catchy slogans, and finally, generated a pretty nice Instagram post to announce the launch of GreenStride!
## A/B testing for marketing comparison
Another interesting field where ChatGPT can assist marketers is A/B testing.
A/B testing in marketing is a method of comparing two different versions of a marketing campaign, advertisement, or website to determine which one performs better. In A/B testing, two variations of the same campaign or element are created, with only one variable changed between the two versions. The goal is to see which version generates more clicks, conversions, or other desired outcomes.
An example of A/B testing might be testing two versions of an email campaign, using different subject lines, or testing two versions of a website landing page, with different call-to-action buttons. By measuring the response rate of each version, marketers can determine which version performs better and make data-driven decisions about which version to use going forward.
A/B testing allows marketers to optimize their campaigns and elements for maximum effectiveness, leading to better results and a higher return on investment.
Since this method involves the process of generating many variations of the same content, the generative power of ChatGPT can definitely assist in that.
Let’s consider the following example. I’m promoting a new product I developed: a new, light and thin climbing harness for speed climbers. I’ve already done some market research and I know my niche audience. I also know that one great channel of communication for that audience is publishing on an online climbing blog, of which most climbing gyms’ members are fellow readers.
My goal is to create an outstanding blog post to share the launch of this new harness, and I want to test two different versions of it in two groups. The blog post I’m about to publish and that I want to be the object of my A/B testing is the following:

Figure 7.12 – An example of a blog post to launch climbing gear
Here, ChatGPT can help us on two levels:
* The first level is that of rewording the article, using different keywords or different attention- grabbing slogans. To do so, once this post is provided as context, we can ask ChatGPT to work on the article and slightly change some elements:

Figure 7.13 – New version of the blog post generated by ChatGPT
As per my request, ChatGPT was able to regenerate only those elements I asked for (title, subtitle, and closing sentence) so that I can monitor the effectiveness of those elements by monitoring the reaction of the two audience groups.
* The second level is working on the design of the web page, namely, changing the collocation of the image rather than the position of the buttons. For this purpose, I created a simple web page for the blog post published in the climbing blog (you can find the code in the book’s GitHub repository at [https://github.com/PacktPublishing/The-Ultimate-Guide- to-ChatGPT-and-OpenAI/tree/main/Chapter%207%20-%20ChatGPT%20 for%20Marketers/Code](https://github.com/PacktPublishing/The-Ultimate-Guide-)):

Figure 7.14 – Sample blog post published on the climbing blog
We can directly feed ChatGPT with the HTML code and ask it to change some layout elements, such as the position of the buttons or their wording. For example, rather than Buy Now, a reader might be more gripped by an I want one! button.
So, lets feed ChatGPT with the HTML source code:

Figure 7.14 – ChatGPT changing HTML code
Let’s see what the output looks like (I also changed the title, subtitle and paragraph with the once generated by ChatGPT):

Figure 7.15 – New version of the website
As you can see, ChatGPT only intervened at the button level, slightly changing their layout, position, color, and wording.
In conclusion, ChatGPT is a valuable tool for A/B testing in marketing. Its ability to quickly generate different versions of the same content can reduce the time to market of new campaigns. By utilizing ChatGPT for A/B testing, you can optimize your marketing strategies and ultimately drive better results for your business.
## Boosting Search Engine Optimization (SEO)
Another promising area for ChatGPT to be a game changer is **Search Engine Optimization** (**SEO**). This is the key element behind ranking in search engines such as Google or Bing and it determines whether your websites will be visible to users who are looking for what you promote.
> SEO is a technique used to enhance the visibility and ranking of a website on **search engine results pages** (**SERPs**). It is done by optimizing the website or web page to increase the amount and quality of organic (unpaid) traffic from search engines. The purpose of SEO is to attract more targeted visitors to the website by optimizing it for specific keywords or phrases.
Imagine you run an e-commerce company called Hat&Gloves, which only sells, as you might have guessed, hats and gloves. You are now creating your e-commerce website and want to optimize its ranking. Let’s ask ChatGPT to list some relevant keywords to embed in our website:

Figure 7.17 – Example of SEO keywords generated by ChatGPT
As you can see, ChatGPT was able to create a list of keywords of different kinds. Some of them are pretty intuitive, such as **Hats** and **Gloves**. Others are related, with an indirect link. For example, **Gift ideas** is not necessarily related to my e-commerce business, however, it could be very smart to include it in my keywords, so that I can widen my audience.
Another key element of SEO is **search engine intent**. Search engine intent, also known as **user intent**, refers to the underlying purpose or goal of a specific search query made by a user in a search engine. Understanding search engine intent is important because it helps businesses and marketers create more targeted and effective content and marketing strategies that align with the searcher’s needs and expectations.
There are generally four types of search engine intent:
* **Informational intent**: The user is looking for information on a particular topic or question, such as What is the capital of France? or How to make a pizza at home.
* **Navigational intent**: The user is looking for a specific website or web page, such as **Facebook login** or **Amazon.com**.
* **Commercial intent**: The user is looking to buy a product or service, but may not have made a final decision yet. Examples of commercial intent searches include *best laptop under $1000* or *discount shoes online*.
* **Transactional intent**: The user has a specific goal to complete a transaction, which might refer to physical purchases or subscribing to services. Examples of transactional intent could be *buy iPhone 13* or *sign up for a gym membership*.
By understanding the intent behind specific search queries, businesses and marketers can create more targeted and effective content that meets the needs and expectations of their target audience. This can lead to higher search engine rankings, more traffic, and ultimately, more conversions and revenue.
Now, the question is, will ChatGPT be able to determine the intent of a given request? Before answering, it is worth noticing that the activity of inferring the intent of a given prompt is the core business of **Large Language Models** (**LLMs**), including GPT. So, for sure, ChatGPT is able to capture prompts’ intents.
The added value here is that we want to see whether ChatGPT is able to determine the intent in a precise domain with a precise taxonomy, that is, the one of marketing. That is the reason why prompt design is once again pivotal in guiding ChatGPT in the right direction.

Figure 7.18 – Example of keywords clustered by user intent by ChatGPT
Finally, we could also go further and leverage once more the Act as… hack, which we already mentioned in *Chapter 3*. It would be very interesting indeed to have an assessment of our website and understand whether it is optimized as intended. In marketing, this analysis is called an **SEO audit**. An SEO audit is an evaluation of a website’s SEO performance and potential areas for improvement. An SEO audit is typically conducted by SEO experts, web developers, or marketers, and involves a comprehensive analysis of a website’s technical infrastructure, content, and backlink profile.
During an SEO audit, the auditor will typically use a range of tools and techniques to identify areas of improvement, such as keyword analysis, website speed analysis, website architecture analysis, and content analysis. The auditor will then generate a report outlining the key issues, opportunities for improvement, and recommended actions to address them.
Let’s ask ChatGPT to act as an SEO expert to conduct this audit. As a reference website, we will use our climbing blog referenced above. I will give ChatGPT the code and ask give the following instructions: “act as a SEO specialist and generate a brief SEO audit (max 300 words) on the above HTML code”. This is the response:

Figure 7.19 – ChatGPT generating an SEO audit on the HTML code of a climbing blog.
ChatGPT was able to generate a pretty accurate analysis, with relevant comments and suggestions. Overall, ChatGPT has interesting potential for SEO-related activities, and it can be a good tool whether you are building your website from scratch, or you want to improve existing ones.
## Sentiment analysis to improve quality and increase customer satisfaction
Sentiment analysis is a technique used in marketing to analyze and interpret the emotions and opinions expressed by customers toward a brand, product, or service. It involves the use of **natural language processing** (**NLP**) and **machine learning** (**ML**) algorithms to identify and classify the sentiment of textual data such as social media posts, customer reviews, and feedback surveys.
By performing sentiment analysis, marketers can gain insights into customer perceptions of their brand, identify areas for improvement, and make data-driven decisions to optimize their marketing strategies. For example, they can track the sentiment of customer reviews to identify which products or services are receiving positive or negative feedback and adjust their marketing messaging accordingly.
Overall, sentiment analysis is a valuable tool for marketers to understand customer sentiment, gauge customer satisfaction, and develop effective marketing campaigns that resonate with their target audience.
Sentiment analysis has been around for a while, so you might be wondering what ChatGPT could bring as added value. Well, besides the accuracy of the analysis (it being the most powerful model on the market right now), ChatGPT differentiates itself from other sentiment analysis tools since it is **general artificial intelligence** (**AGI**).
This means that when we use ChatGPT for sentiment analysis, we are not using one of its specific APIs for that task: the core idea behind ChatGPT and OpenAI models is that they can assist the user in many general tasks at once, interacting with a task and changing the scope of the analysis according to the user’s request.
So, for sure, ChatGPT is able to capture the sentiment of a given text, such as a Twitter post or a product review. However, ChatGPT can also go further and assist in identifying specific aspects of a product or brand that are positively or negatively impacting the sentiment. For example, if customers consistently mention a particular feature of a product in a negative way, ChatGPT can highlight that feature as an area for improvement. Or, ChatGPT might be asked to generate a response to a particularly delicate review, keeping in mind the sentiment of the review and using it as context for the response. Again, it can generate reports that summarize all the negative and positive elements found in reviews or comments and cluster them into categories.
Let’s consider the following example. A customer has recently purchased a pair of shoes from my e-commerce company, RunFast, and left the following review:
*“I recently bought the RunFast Prodigy shoes and have mixed feelings. They're extremely comfortable with excellent cushioning and support, reducing foot fatigue during my runs. The design is also appealing, and I've received several compliments. However, the durability is disappointing; the outsole wears quickly, and the breathable upper shows signs of wear after a few weeks. Given the high price, I'm hesitant to recommend them despite their comfort and design.”*
Let’s ask ChatGPT to capture the sentiment of this review:

Figure 7.20 – ChatGPT analyzing a customer review
From the preceding figure, we can see how ChatGPT didn’t limit itself to providing a label: it also explained both the positive and negative elements characterizing the review, which has a mixed feeling and hence can be labeled as neutral overall.
Let’s try to go deeper into that and ask some suggestions about improving the product:

Figure 7.21 – Suggestions on how to improve my product based on customer feedback
Finally, let’s generate a response to the customer, showing that we, as a company, do care about customers’ feedback and want to improve our products.

Figure 7.22 – Response generated by ChatGPT
The example we saw was a very simple one with just one review. Now imagine we have tons of reviews, as well as diverse sales channels where we receive feedback. Imagine the power of tools such as ChatGPT and OpenAI models, which are able to analyze and integrate all of that information and identify the pluses and minuses of your products, as well as capturing customer trends and shopping habits. Additionally, for customer care and retention, we could also automate review responses using the writing style we prefer. In fact, by tailoring your chatbot’s language and tone to meet the specific needs and expectations of your customers, you can create a more engaging and effective customer experience.
Here are some examples:
* Empathetic chatbot: A chatbot that uses an empathetic tone and language to interact with customers who may be experiencing a problem or need help with a sensitive issue
* Professional chatbot: A chatbot that uses a professional tone and language to interact with customers who may be looking for specific information or need help with a technical issue
* Conversational chatbot: A chatbot that uses a casual and friendly tone to interact with customers who may be looking for a personalized experience or have a more general inquiry
* Humorous chatbot: A chatbot that uses humor and witty language to interact with customers who may be looking for a light-hearted experience or to diffuse a tense situation
* Educational chatbot: A chatbot that uses a teaching style of communication to interact with customers who may be looking to learn more about a product or service
In conclusion, ChatGPT can be a powerful tool for businesses to conduct sentiment analysis, improve their quality, and retain their customers. With its advanced natural language processing capabilities, ChatGPT can accurately analyze customer feedback and reviews in real time, providing businesses with valuable insights into customer sentiment and preferences. By using ChatGPT as part of their customer experience strategy, businesses can quickly identify any issues that may be negatively impacting customer satisfaction and take corrective action. Not only can this help businesses improve their quality, but it can also increase customer loyalty and retention.
## Summary
In this chapter, we explored ways in which ChatGPT can be used by marketers to enhance their marketing strategies. We learned that ChatGPT can help in developing new products as well as defining their go-to-market strategy, designing A/B testing, enhancing SEO analysis, and capturing the sentiment of reviews, social media posts, and other customer feedback.
The importance of ChatGPT for marketers lies in its potential to revolutionize the way companies engage with their customers. By leveraging the power of NLP, ML, and big data, ChatGPT allows companies to create more personalized and relevant marketing messages, improve customer support and satisfaction, and ultimately, drive sales and revenue.
As ChatGPT continues to advance and evolve, it is likely that we will see even more involvement in the marketing industry, especially in the way companies engage with their customers. In fact, relying heavily on AI allows companies to gain deeper insights into customer behavior and preferences.
The key takeaway for marketers is to embrace these changes and adapt to the new reality of AI-powered marketing in order to stay ahead of the competition and meet the needs of their customers.
In the next chapter, we will look at the third and last domain in the application of ChatGPT covered in this book – research.
# 7 Research Reinvented with ChatGPT
## Join our book community on Discord

[https://packt.link/EarlyAccess](https://packt.link/EarlyAccess)
In this chapter, we focus on researchers who wish to leverage ChatGPT. The chapter will go through a few main use cases ChatGPT can address, so that you will learn from concrete examples how ChatGPT can be used in research.
By the end of this chapter, you will be familiar with using ChatGPT as a research assistant in many ways, including the following:
* Researchers’ need for ChatGPT
* Brainstorming literature for your study
* Providing support for the design and framework of your experiment
* Generating and formatting the bibliography to incorporate in your research study
Delivering a pitch or slide deck presentation about your study addressing diverse audiences This chapter will also provide examples and enable you to try the prompts on your own.
## Researchers’ need for ChatGPT
ChatGPT can be an incredibly valuable resource for researchers across a wide range of fields. As a sophisticated language model trained on vast amounts of data, ChatGPT can quickly and accurately process large amounts of information and generate insights that might be difficult or time-consuming to uncover through traditional research methods.
Additionally, ChatGPT can provide researchers with a unique perspective on their field, by analyzing patterns and trends that might not be immediately apparent to human researchers. For example, imagine a researcher studying climate change and wanting to understand the public perception of this issue. They might ask ChatGPT to analyze social media data related to climate change and identify the most common themes and sentiments expressed by people online. ChatGPT could then provide the
researcher with a comprehensive report detailing the most common words, phrases, and emotions associated with this topic, as well as any emerging trends or patterns that might be useful to know.
By working with ChatGPT, researchers can gain access to cutting-edge technology and insights, and stay at the forefront of their field.
Let’s now dive deeper into four use cases where ChatGPT can boost research productivity.
> Most examples proposed in this Chapter are based on up-to-date information; in fact, you will see ChatGPT leveraging the web search plug-in very often. If you are using ChatGPT with the GPT-3.5-turbo (so, the free edition), keep in mind that it doesn’t have the web search plug-in enabled and, henceforth, its knowledge cutoff is limited to the 2021\. This could be a limitation if you are looking for updated information and, in general, for references coming from the web (so that you can double-check the groundedness of the response).
## Brainstorming literature for your study
A literature review is a critical and systematic process of examining existing published research on a specific topic or question. It involves searching, reviewing, and synthesizing relevant published studies and other sources, such as books, conference proceedings, and gray literature. The goal of a literature review is to identify gaps, inconsistencies, and opportunities for further research in a particular field.
The literature review process typically involves the following steps:
1. **Defining the research question**: The first step in conducting a literature review is to define the research question of the topic of interest. So, let’s say we are carrying on a research on the effects of social media on mental health. Now we are interested in brainstorming some possible research questions to focus our research on, and we can leverage ChatGPT to do so:

Figure 8.1 – Examples of research questions based on a given topic
Those are all interesting questions that could be further investigated. Since I’m particularly interested in the first one – “In what ways do social media interactions and online support communities affect the mental well-being of individuals with chronic mental health conditions?” - I will keep that one as a reference for the next steps of our analysis.
1. **Searching for literature**: Now that we have our research question, the next step is to search for relevant literature using a variety of databases, search engines, and other sources. Researchers can use specific keywords and search terms to help identify relevant studies.

Figure 8.2 – Literature search with the support of ChatGPT
Starting from the suggestions of ChatGPT, we can start diving deeper into those references.
1. **Screening the literature**: Once relevant literature has been identified, the next step is to screen the studies to determine whether they meet the inclusion criteria for the review. This typically involves reviewing the abstract and, if necessary, the full text of the study.
Let’s say, for example, that we want to go deeper into the **Social Media and Teen Anxiety: Insights from Harvard Graduate School of Education** research paper. Let’s ask ChatGPT to screen it for us:

Figure 8.3 – Literature screening of a specific paper
ChatGPT was able to provide me with an overview of the paper and, considering its research question and main topics of discussion, I think it will be pretty useful for my own study.
1. **Extracting data**: After the relevant studies have been identified, researchers will need to extract data from each study, such as the study design, sample size, data collection methods, and key findings.
For example, let’s say that we want to gather the following information from the paper Digital Self-Harm: Prevalence, Motivations, and Outcomes by Hinduja and Patchin (2018):
* Data sources collected in the paper and object of the study
* data collection method adopted by researchers
* data sample size
* main limitations and drawbacks of the analysis
* the experiment design adopted by researchers
Here is how it goes

Figure 8.4 – Extracting relevant data and frameworks from a given paper
1. **Synthesizing the literature**: The final step in the literature review process is to synthesize the findings of the studies and draw conclusions about the current state of knowledge in the field. This may involve identifying common themes, highlighting gaps or inconsistencies in the literature, and identifying opportunities for future research.
Let’s imagine that, besides the papers proposed by ChatGPT, we have collected other titles and papers we want to synthesize. More specifically, I want to understand whether they drive the same conclusions, what are the common trends, and which method might be more reliable than others. For this scenario, we will consider three research papers:
* *The Effects of Social Media on Mental Health: A Proposed Study*, by Grant Sean Bossard ([https://digitalcommons.bard.edu/cgi/viewcontent. cgi?article=1028&context=senproj_f2020](https://digitalcommons.bard.edu/cgi/viewcontent.))
* *The Impact of Social Media on Mental Health*, by Vardanush Palyan ([https://www. spotlightonresearch.com/mental-health-research/the-impact-of- social-media-on-mental-health](https://www.))
* *The Impact of Social Media on Mental Health: a mixed methods research of service providers’ awareness*, by Sarah Nichole Koehler and Bobbie Rose Parrell ([https://scholarworks. lib.csusb.edu/cgi/viewcontent.cgi?article=2131&context=etd](https://scholarworks.))
Here is how the results appear:

Figure 8.5 – Literature analysis and benchmarking of three research papers
Also, in this case, ChatGPT was able to produce a relevant summary and analysis of the three papers provided, including benchmarking among the methods and reliability considerations.
Overall, ChatGPT was able to carry out many activities in the field of literature review, from research question brainstorming to literature synthesis. As always, a **subject-matter expert** (**SME**) is needed in the loop to review the results, however, with this assistance many activities can be done more efficiently.
Another activity that can be supported by ChatGPT is the design of the experiment the researcher wants to carry out. We are going to look at that in the following section.
## Providing support for the design and framework of your experiment
Experiment design is the process of planning and executing a scientific experiment or study to answer a research question. It involves making decisions about the study’s design, the variables to be measured, the sample size, and the procedures for collecting and analyzing data.
ChatGPT can help in experiment design for research by suggesting to you the study framework, such as a randomized controlled trial, quasi-experimental design, or a correlational study, and supporting you alongside the implementation of that design.
Let’s consider the following scenario. We want to investigate the effects of a new educational program on student learning outcomes in mathematics. This new program entails **project-based learning** (**PBL**), meaning that students are asked to work collaboratively on real-world projects, using math concepts and skills to solve problems and create solutions.
For this purpose, we defined our research question as follows:
How does the new PBL program compare to traditional teaching methods in improving student performance?
Here’s how ChatGPT can help:
* **Determining study design**: ChatGPT can assist in determining the appropriate study design for the research question, such as a randomized controlled trial, quasi-experimental design, or correlational study.

Figure 8.6 – ChatGPT suggesting the appropriate study design for your experiment
ChatGPT suggested proceeding with a randomized controlled trial (RCT) and provided a clear explanation of the reason behind it. It seems reasonable to me to proceed with this approach: the next steps will be to identify outcome measures and variables to consider in our experiment.
* **Identifying outcome measures**: ChatGPT can help you identify some potential outcome measures to determine the results of your test. Let’s ask for some suggestions for our study:

Figure 8.7 – Learning outcomes for the given research study
It is reasonable for me to pick test scores as the outcome measure.
* **Identifying variables**: ChatGPT can help the researcher to identify the independent and dependent variables in the study:

Figure 8.8 – ChatGPT generating variables for the given study
Note that ChatGPT was also able to generate the type of variables, called **control variables**, that are specific to the study design we are considering (RCT).
Control variables, also known as covariates, are variables that are held constant or are controlled in a research study in order to isolate the relationship between the independent variable(s) and the dependent variable. These variables are not the primary focus of the study but are included to minimize the effect of confounding variables on the results. By controlling these variables, researchers can reduce the risk of obtaining false positive or false negative results and increase the internal validity of their study.
With the preceding variables, we are ready to set up our experiment. Now we need to select participants, and ChatGPT can assist us with that.
* **Sampling strategy**: ChatGPT can suggest potential sampling strategies for the study:

Figure 8.9 – RCT sampling strategy suggestion from ChatGPT
Note that, in a real-world scenario, it is always a good practice to ask AI tools to generate more options with explanations behind them, so that you can make a reasoned decision. For this example, let’s go ahead with what ChatGPT suggested to us, which also includes suggestions about the population of interest and sample size.
* **Data analysis**: ChatGPT can assist the researcher in determining the appropriate statistical tests to analyze the data collected from the study, such as ANOVA, t-tests, or regression analysis.

Figure 8.10 – ChatGPT suggests a statistical test for a given study
Everything suggested by ChatGPT is coherent and finds confirmation in papers about how to conduct a statistical test. It was also able to identify that we are probably talking about a continuous variable (that is, scores) so that we know that all the information ahead is based on this assumption. In case we want to have discrete scores, we might adjust the prompt by adding this information, and ChatGPT will then suggest a different approach.
The fact that ChatGPT specifies assumptions and explains its reasoning is key to making safe decisions based on its input.
In conclusion, ChatGPT can be a valuable tool for researchers when designing experiments. By utilizing its **natural language processing** (**NLP**) capabilities and vast knowledge base, ChatGPT can help researchers select appropriate study designs, determine sampling techniques, identify variables and learning outcomes, and even suggest statistical tests to analyze the data.
In the next section, we are going to move forward in exploring how ChatGPT can support researchers, focusing on bibliography generation.
## Generating and formatting a bibliography
ChatGPT can support researchers in bibliography generation by providing automated citation and reference tools. These tools can generate accurate citations and references for a wide range of sources, including books, articles, websites, and more. ChatGPT knows various citation styles, such as APA, MLA, Chicago, and Harvard, allowing researchers to select the appropriate style for their work. Additionally, ChatGPT can also suggest relevant sources based on the researcher’s input, helping to streamline the research process and ensure that all necessary sources are included in the bibliography. By utilizing these tools, researchers can save time and ensure that their bibliography is accurate and comprehensive.
Let’s consider the following example. Let’s say we finalized a research paper titled, The Impact of Technology on Workplace Productivity: An Empirical Study. During the research and writing process, we collected the following references to papers, websites, videos, and other sources that we need to include in the bibliography (in order, three research papers, one YouTube video, and one website):
* *The second machine age: Work, progress, and prosperity in a time of brilliant technologies*. Brynjolfsson, 2014. https://psycnet.apa.org/record/2014-07087-000
* *The Impact of Technostress on Role Stress and Productivity. Tarafdar*, 2014\. Pages 301-328. https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109
* *The big debate about the future of work, explained*”. Vox. https://www.youtube.com/ watch?v=TUmyygCMMGA
Obviously, we cannot have the preceding list in our research paper; we need proper formatting for it. To do so, we can provide ChatGPT with the raw list of references and ask it to re-generate it with the specific format, for example, APA style, the official style of the **American Psychological Association** (**APA**), commonly used as reference format style in education, psychology, and social sciences.
Let’s see how ChatGPT works with that:

Figure 8.11 – A list of references generated in APA format by ChatGPT
Note that I specified not to add details in case ChatGPT doesn’t know them. Indeed, I noticed that sometimes ChatGPT was adding the month and day of publication, making some mistakes.
Other interesting assistance ChatGPT can provide is that of suggesting potential reference papers we might want to quote. We’ve already seen in the first paragraph of this chapter how ChatGPT is able to brainstorm relevant literature before the writing process; however, once the paper is done, we might have forgotten to quote relevant literature, or even not be aware of having quoted someone else’s work.
ChatGPT can be a great assistant in brainstorming possible references we might have missed. Let’s consider once more our paper, focused on the research question “In what ways do social media interactions and online support communities affect the mental well-being of individuals with chronic mental health conditions?”. Let’s say that we set the following title and abstracts: which has the following abstract:
**Title***: The Impact of Social Media and Online Support Communities on Mental Well-being in Individuals with Chronic Mental Health Conditions*
**Abstract***: This study explores the effects of social media interactions and online support communities on the mental well-being of individuals with chronic mental health conditions. By examining various online platforms, the research aims to identify both positive and negative impacts. The study utilizes qualitative and quantitative methods, including surveys and interviews, to assess changes in anxiety, depression, and overall life satisfaction. Preliminary findings suggest that while online support can offer valuable social connections and emotional support, excessive use and negative interactions may exacerbate mental health issues.*
Let’s ask ChatGPT to list all the possible references that might be related to this kind of research:

Figure 8.12 – List of references related to the provided abstract
You can repeat this process with other sections of your paper also, to make sure you are not missing any relevant references to include in your bibliography.
Once you have your study ready, you will probably need to present it with an elevator pitch. In the next section, we will see how ChatGPT can also support this task.
## Generating a presentation of the study
The last mile of a research study is often that of being presented to various audiences. This might involve preparing a slide deck, pitch, or webinar where the researcher needs to address different kinds of audiences.
Let’s say, for example, that our study *The Impact of Social Media and Online Support Communities on Mental Well-being in Individuals with Chronic Mental Health Conditions* is meant for a master’s degree thesis discussion. In that case, we can ask ChatGPT to produce a pitch structure that is meant to last 15 minutes and adheres to the scientific method. Let’s see what kind of results are produced (as context, I’m referring to the abstract of the previous paragraph):

Figure 8.13 – Thesis discussion generated by ChatGPT
That was impressive! Back in my university days, it would have been useful to have such a tool to assist me in my discussion design.
Starting from this structure, we can also ask ChatGPT to generate a slide deck as a visual for our thesis discussion.
Let’s proceed with this request:

Figure 8.14 – Slide deck structure based on a discussion pitch
Finally, let’s imagine that our thesis discussion was outstanding to the point that it might be selected for receiving research funds in order to keep investigating the topic. Now we need an elevator pitch to convince the funding committee. Let’s ask for some support from ChatGPT:

Figure 8.15 – Elevator pitch for the given thesis
We can always adjust results and make them more aligned to what we are looking for, however, having structures and frameworks already available can save a lot of time and allows us to focus more on the technical content we want to bring.
Overall, ChatGPT is able to support an end-to-end journey in research, from literature collection and review to the generation of the final pitch of the study, and we’ve demonstrated how it can be a great AI assistant for researchers.
Furthermore, note that in the field of research, some tools that are different from ChatGPT, yet still powered by GPT models, have been developed recently. An example is Humanata.ai, an AI-powered tool that allows you to upload your documents and perform several actions on them, including summarization, instant Q&A, and new paper generation based on uploaded files.
This suggests how GPT-powered tools (including ChatGPT) are paving the way toward several innovations within the research domain.
## Summary
In this chapter, we explored the use of ChatGPT as a valuable tool for researchers. Through literature review, experiment design, bibliography generation and formatting, and presentation generation, ChatGPT can assist the researcher in speeding up those activities with low or zero added value, so that they can focus on relevant activities.
Note that we focused only on a small set of activities where ChatGPT can support researchers. There are many other activities within the domain of research that could benefit from the support of ChatGPT, among which we can mention data collection, study participant recruitment, research networking, public engagement, and many others.
Researchers who incorporate this tool into their work can benefit from its versatility and time-saving features, ultimately leading to more impactful research outcomes.
However, it is important to keep in mind that ChatGPT is only a tool and should be used in conjunction with expert knowledge and judgment. As with any research project, careful consideration of the research question and study design is necessary to ensure the validity and reliability of the results.
With this chapter, we also close Part 2 of this book, which focused on the wide range of scenarios and domains you can leverage ChatGPT for. However, we mainly focused on individual or small team usage, from personal productivity to research assistance. Starting from Part 3, we will elevate the conversation to how large organizations can leverage the same generative AI behind ChatGPT for enterprise-scale projects, using OpenAI model APIs available on the Microsoft Azure cloud.
## References
* [https://arxiv.org/abs/2312.11914](https://arxiv.org/abs/2312.11914)
* [https://arxiv.org/abs/2101.07714](https://arxiv.org/abs/2101.07714)
* [https://arxiv.org/abs/2101.07714](https://arxiv.org/abs/2101.07714)
* [https://psycnet.apa.org/record/2014-07087-000](https://psycnet.apa.org/record/2014-07087-000)
* [https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109](https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109)
* [https://www.youtube.com/watch?v=TUmyygCMMGA](https://www.youtube.com/watch?v=TUmyygCMMGA)
# 9 Exploring GPTs
## Join our book community on Discord

[https://packt.link/EarlyAccess](https://packt.link/EarlyAccess)
In last Chapters, we saw several examples of how to leverage ChaGPT for various activities, from personal productivity to marketing, from research to software development. For each of these scenarios, we always faced a similar situation: we started with a general purpose model as ChatGPT, to then ask very specific questions or providing specific references to tailor it to our specific need.
However, sometimes this might not be enough if our aim is that of obtaining an extremely specialized model for our own purposes. That’s why we might need to build “purpose-specific ChatGPT” and, luckily enough, OpenAI itself developed a no-code platform to build those customized assistants, that are called GPTs.
In this Chapter, we are going through the GPTs features, capabilities and real-world applications, covering the same use cases we saw in previous Chapters, so that you can see the difference in the output’s quality. Plus, we will also see how to publish your GPT and make it a production application not only for yourself but also for company or for everyone else.
By the end of this Chapter, you will be able to:
* Understand what a GPT is and what kind of tasks it can achieve
* Build your own GPT without writing a line of code
* Publish your GPT and integrate it with external systems
Let’s start from some basic definitions and then jump into the practice.
## Technical requirements
ChatGPT Plus
## What are GPTs?
In November 2023, OpenAI introduced GPTs, specialized versions of ChatGPT designed to enhance productivity and cater to specific tasks and needs. Unlike the general-purpose ChatGPT, these custom versions, called GPTs, allow users to create tailored AI models without any coding knowledge.
> It is important to anticipate a taxonomy consideration that will be relevant throughout this chapter. When we mention the word GPT, there are two main definitions that you found - and will find – in this book:
* The first one refers to the proper “Generative Pre-trained Transformer” (GPT) model architecture behind OpenAI’s language models. We already mentioned this architecture in Part 1, and we know that this is the framework behind ChatGPT itself.
* The second one refers, in a more generic way, to the specialized assistants that OpenAI allows users to create in a no-code approach. With GPTs, OpenAI refers to specialized versions of ChatGPT and, in this context, a single GPT refers to one assistant that you created leveraging the GPTs platform (available to all users with ChatGPT Plus).
Over this chapter, whenever you will read GPT or GPTs, keep in mind that we are using the second definition.
The idea of GPTs is like that of AI Agents. In fact, with GPTs, we are building entities powered by LLMs, with specific instructions and provided with both custom knowledge base and a set of tools or plug-ins to interact with the surrounding environment.
Let’s see more in details all these components. First of all, you can have a comprehensive look at all existing GPTs that have been publicly published. To do so, you can navigate through the following page: [https://chatgpt.com/gpts](https://chatgpt.com/gpts).

Figure 1: Landing page of GPTs.
As you can see from the preceding picture, this is a proper marketplace of all existing GPTs. You can explore it by category (box 1) or by explaining what you are looking for in the search bar (box 2).
Then, to create your own GPT, you can navigate through the https://chatgpt.com/gpts/editor page and login with your OpenAI account. Alternatively, you can go to ChatGPT and select on the top-left corner of your page “Explore GPTs”, then “Create” (as illustrated in box 3 in the preceding picture).
Once in the editor, you will be asked to configure your GPT, while on the right-hand side you have the possibility to test it in real time.

Figure 2: Editor page to create your GPT from scratch.
Let’s explore all the components:
* **Name**: the name you give to your GPT.
* **Description**: the description of the capabilities of your GPT. This is extremely important especially if you are going to publish your GPT in the marketplace, so that other users can find it easily (as we previously mentioned, GPTs can be searched via natural language in the Search GPTs bar).
* **Instructions**: this is the metaprompt of your GPT, so that set of instructions in natural language that tailor the assistant to your specific need and that the final user doesn’t see.
* **Conversation Starters**: this is a set of sample prompts that users could start with to interact with the GPT and gain confidence.
* **Knowledge**: this refers to the custom documentation we can ground the model on. When we upload documents here, our GPT will be able to navigate through it with the Retrieval Augmented Generation pattern, so that we can provide additional knowledge or even limit our assistant’s responses to the custom knowledge base, depending on our needs.
* **Capabilities**: these refer to a set of built-in plug-ins we can provide the GPT with, without writing a single line of code. You can see from the above picture that there are three plug-ins out of the box:
* Web browsing for searching the web and retrieve up to date information.
* DALL-E 3 to generate illustrations.
* Code Interpreter and Data Analysis to execute code in a sandboxed Python environment and interact with analytical files (like spreadsheets).
* Actions: actions can be seen as plug-ins as well, however they differ from capabilities since they are not built-in, but rather specified by the GPT developer. For example, you might generate an action
> In the context of actions, you can be supported by a specialized GPT which is natively integrated in the configuration pane if you click on “Add Actions” and then “Get help from ActionsGPT”.

Figure 3: Example on how to get support from ActionsGPT
By doing so, you will be prompted to the chat interface of your GPT:

Figure 4: Landing page of ActionsGPT.
It’s amazing to see how we are witnessing “GPTs inside GPTs” approaches, don’t you think?
In addition to the standard configuration page, there is another option to build your GPT in a more “conversational way”. In fact, you can switch to the tab “Create” and explain in natural language what you want to achieve with your GPT:

Figure 5: Creating a GPT starting from a conversation in natural language.
We will see both methods – standard configuration and conversational configuration – in the practical paragraphs of this chapter.
Now that we know what a GPT is, let’s see how to create one. Over the next sections, we will create 5 different GPTs:
* The first four will specialize in the four domains we already covered with the general-purpose ChatGPT – personal productivity, code development, marketing and research. The idea is that to compare the overall efficiency and accuracy of a domain-specific GPT versus a general-purpose ChatGPT.
* The fifth one will cover a new domain, related to artistic creativity. We will leverage the built-in DALL-E plug-in as well as other plug-ins developed by design companies (like Canvas).
Let’s start with boosting our personal productivity with a specialized GPT.
## Personal Assistant
In this scenario, we are going to build a GPT to boost our gym workouts. To do so, we will leverage the built-in plug-ins. Plus, we will add some relevant documents so that the assistant will be grounded on specific knowledge base.
My goal is that of having an assistant that can build for me workout plans according to by fitness ambition, availability, gender and age, preferences and so forth. I also want my assistant to provide me clear explanations on the reason behind its proposals.
To accomplish all of this, I want to make sure that my assistant will:
* Ask me specific questions it needs to design the best workout for me
* Provide me with relevant information and sources regarding its responses
* Consider my feedback, yet it will be able to maintain its idea if it believes it is correct for me.
* Not accommodate my requests if not reasonable or risky for my health.
Let’s see how to create our GPT, following all the configuration steps mentioned in the previous section:
1. **Name**: I’ll call my assistant WorkoutGPT.
2. **Description**: this is the description I set: “Workout Assistant that helps user designing their workout plan according to their needs.”
3. **Instructions**: here we start with the real core of our GPT. This is the set of instructions I provided the GPT with:
*“You are a workout AI assistant that help users creating their workout plans, depending on their needs.*
*Before generating the plan, make sure to ask the following questions:*
*- Fitness goal and time expectation*
*- Age and gender*
*- Fitness level*
*- Availability for the workout*
*- All other elements that you need in order to define a proper workout plan (e.g. equipment, potential injuries...)*
*If needed, use the provided documents to enrich your responses.*
*If the user suggests you something that it's not plausible for its goal, remain stick to your assumption, explaining the reason behind politely.*
*If the user asks you something that might be risky for his/her health, politely suggest to take it easier and an alternative approach, explaining the reason behind.”*
1. **Conversation Starters**: here I set three samples of different workouts:
* I want to train for a marathon in 6 months.
* Generate a 30’ HIIT workout plan without any equipment.
* Generate a 45’ weightlifting workout plan with dumbbells only.
2. **Knowledge**: here I uploaded the standard National Strength and Conditioning Association (NSCA) training load chart, a tool used to help athletes and coaches determine the appropriate training load for different exercises and training sessions. It looks like following:

Figure 6: NSCA training load chart. Source: https://www.nsca.com/contentassets/61d813865e264c6e852cadfe247eae52/nsca_training_load_chart.pdf
I almost forgot a key step – the illustration! It might seem superficial yet having an icon for your GPT make it more attractive, especially if you are planning to publish it to all users. Luckily enough, we have DALL-E directly integrated in the configuration pane:

Figure 7: Example on how to set your GPT icon.
This is how the configuration looks like:

Figure 8: WorkoutGPT configuration page.
Great! Let’s now see some sample conversations.
To start with, let’s pick the first conversation starter about the marathon training:

Figure 9: Example of questions asked from WorkoutGPT to assess user’s overall goals and fitness level.
As you can see, our WorkoutGPT immediately asks us the required information to proceed with the plan. Once provided the above information, my assistant generated the following 24-weeks plan (I’ll share here the first 4 weeks):

Figure 10: Example of workout table for marathon training generated by WorkoutGPT.
Plus, it specifies the Friday strength training as follows:

Figure 11: Example of strength training generated by WorkoutGPT.
Let’s focus on the strength training. I want to better understand how to calibrate the weights. The following illustration shows the first extract of the response:

Figure 12: Example of WorkoutGPT explaining how to determine the weight to lift in a strength workout.
In the same response, the Assistant also referenced the NCSA training load chart provided as knowledge base:

Figure 13: Example of WorkoutGPT retrieving information from custom knowledge base about NSCA training load chart.
Now I want to challenge my WorkoutGPT, asking for something that might be harmful for me. For example, preparing a marathon with no experience in one month is definitely a terrible idea. Let’s see what my assistant’s thoughts about it are, once I provide it with the list of answers to its starter questions:

Figure 14: Example of WorkoutGPT gently nudging the user to pivot its goal and expectations given the risks associated with the request.
As you can see, my WorkoutGPT is nudging me to change my approach to the race. While it still provides me with a running workout plan for 3 weeks (here the output is truncated), it is not aimed at running a marathon under 3h and 15m, but rather on building endurance and strength.
Note that, if we had asked the same thing to the general purpose ChatGPT, it would have responded as follows:

Figure 15: Example of ChatGPT accomplishing user’s task despite the associated risk.
Note how ChatGPT, despite being vocal in disclosing its concerns, it's still accommodating my request, providing me with a plan to run a full marathon. This could encourage me – a reckless beginner who thinks that running a marathon is a joke – to dive into this dumb adventure, with serious consequences for me health.
Overall, GPTs allow you to be extremely specific about how your assistant should behave and what it should avoid saying or accommodate.
## Code Assistant
In this section, I want to develop an assistant that is tailored towards data science projects. More specifically, I want my assistant to be able to:
* Provide clear guidance on how to set up a data science experiment, given the user’s task
* Generate the Python code needed to run the experiment
* Leverage the code interpreter capabilities to run and check the code
* Push the final code to the GitHub repository.
Let’s see how to do that step by step.
1. Setting the instructions:
*This GPT is a data science assistant that helps users set up and run data science experiments. It provides clear guidance on how to define and organize tasks, generate Python code needed for the experiments, and leverages code interpreter capabilities to run and check the code. The GPT will take the user's input and provide step-by-step instructions to structure the experiment, create the necessary scripts, and execute the code.*
*The GPT will execute the code to see whether it works. Once the final code is accepted by the user, it can be pushed to the GitHub repo as a .ipynb file.*
1. Setting the (optional) conversation starters:
*How do I set up a classification experiment?*
*Generate Python code for a random forest model.*
*Can you help me preprocess this dataset?*
*Run this code and check for errors.*
1. Enabling the Code Interpreter & Data Analysis plugin:

Figure 16: Enabling the code Interpreter & Data analysis plugin.
2. Creating an Action to communicate with GitHub. To do so, we need to click on “Create new action” and define the required Schema.
> To set the schema of a ChatGPT action using an OpenAPI 3.1.0 specification, you define the structure of the data (requests and responses) that the action will handle. This involves specifying the schema property under content in your API paths. The schema outlines the expected data types, required fields, and possible values for your request and response bodies.
Steps to Define a Schema
1. Identify the Data Structure: Determine the type of data the action will handle (e.g., JSON objects, arrays).
2. Define Properties: Under the schema, specify the properties, their types, and any constraints. For example, if the action requires a username and an email, you'd define these under properties.
3. Set Required Fields: Use the required array to specify which fields must be provided.
4. Apply to Paths: Place the schema under the appropriate HTTP method in your paths (e.g., POST, GET).
Let’s consider the following example:
openapi: 3.1.0
info:
title: ChatGPT Action API
version: 1.0.0
paths:
/perform-action:
post:
operationId: performAction
summary: Perform a specific action with given inputs.
requestBody:
required: true
content:
application/json:
schema:
type: object
properties:
action:
type: string
description: The action to perform
parameters:
type: object
description: Parameters for the action
properties:
userId:
type: string
content:
type: string
required:
- action
- parameters
responses:
'200':
description: Successful action response
content:
application/json:
schema:
type: object
properties:
success:
type: boolean
message:
type: string
In this example, the `requestBody` for the POST method defines a schema with action and parameters as required fields. The response is also defined, specifying the structure of the data returned after the action is performed.
This is how the schema looks like:

Figure 17: Configuration of the GitHub Action schema.
This is the full schema I used:
openapi: 3.1.0
info:
title: GitHub API
description: API for interacting with GitHub, including pushing code to a repository.
version: 1.0.0
servers:
- url: https://api.github.com
description: GitHub API server
paths:
/repos/{owner}/{repository}/contents/{path}:
put:
operationId: updateFileContents
summary: Create or update a file in a GitHub repository
description: Use this endpoint to create a new file or update an existing file in a repository.
parameters:
- name: owner
in: path
required: true
description: The owner of the repository.
schema:
type: string
- name: repo
in: path
required: true
description: The name of the repository.
schema:
type: string
- name: path
in: path
required: true
description: The file path in the repository.
schema:
type: string
requestBody:
required: true
content:
application/json:
schema:
type: object
properties:
message:
type: string
description: Commit message for the file change.
content:
type: string
description: The new file content, Base64 encoded.
sha:
type: string
description: SHA of the file being replaced, if updating.
branch:
type: string
description: The branch where the file should be created or updated.
committer:
type: object
properties:
name:
type: string
email:
type: string
required:
- message
- content
responses:
'200':
description: Successful file update or creation.
'201':
description: Successful file creation.
'422':
description: Validation failed or the file already exists.
security:
- bearerAuth: []
components:
securitySchemes:
bearerAuth:
type: http
scheme: bearer
bearerFormat: token
schemas: # This subsection contains schema definitions (if needed).
ExampleSchema: # Example object schema
type: object
properties:
exampleProperty:
type: string
To allow the Action to communicate with my repo, I created an access token in my GitHub profile under Settings>Developer Settings>Personal access token.
You can then test the connection as follows:

Figure 18: Example of testing the action.
Let’s now navigate through the Repo and see whether it worked:

Figure 19: Example of file uploaded via the GPT Action.
Great! As you can see, we now have a new file with the pre-defined content.
Let’s now create it and test it:
* Let’s start with a simple question about how to set up a classification experiment (truncated output):

Figure 20: Example of the DataScience Assistant providing guidance on how to tackle a classification problem.
* Following these instructions, we are now ready to set up our experiment. We want to tackle the well-known Titanic passengers’ survival prediction. We will upload our dataset (you can find many free versions online, I downloaded mine at [https://github.com/datasciencedojo/datasets/blob/master/titanic.csv](https://github.com/datasciencedojo/datasets/blob/master/titanic.csv)) and leverage a Logistic Regression model.
> The Titanic survival prediction task is a classic problem in data science and machine learning. The goal is to predict whether a passenger on the Titanic would survive or not, based on various features like age, gender, passenger class, fare, and more. This task is typically used to teach classification techniques, where the model is trained on a labeled dataset and then used to predict outcomes on new data. The challenge involves selecting relevant features, handling missing data, and choosing the appropriate machine learning algorithm to achieve accurate predictions.
* This is my query:

Figure 21: Example of DataScience Assistant designing the Titanic survival experiment.
I will now share some screenshots of the model’s response per each step:
1. Load the data:

Figure 22: Example of DataScience Assistant executing the step 1 of the experiment.
2. Explore and clean the data:

Figure 23: Example of DataScience Assistant exploring and cleaning data.
Note that, when you see the symbol [**>_**], it means that the code interpreter plugin has been triggered. You can click on it to see the executed code:

Figure 24: Example of DataScience Assistant’s generated code leveraging the code interpreter plugin.
1. Feature Engineering:

Figure 25: Example of DataScience Assistant doing feature engineering.
2. Data splitting:

Figure 26: Example of DataScience Assistant splitting the dataset into training and test set.
3. Train the model:

Figure 27: Example of DataScience Assistant training a Logistic Regression model.
4. Evaluate the model:

Figure 28: Example of DataScience Assistant evaluating the model’s output.
That’s pretty cool! It is extremely accurate and can save a lot of time. Plus, if you think about data scientists working in large enterprises on many projects, having a similar assistant can also help in following fixed standards among projects, so that the maintenance is aligned across teams.
The last thing that we will ask our GPT is to push the code on our Repo. Let’s see how it works:

Figure 29 Example of DataScience Assistant leveraging the Action to push the code on GitHub.
If we click on the link, we can see that the file has been successfully uploaded:

Figure 30: File uploaded via DataScience Assistant Action.
And it worked! Again, this is an example of how GPTs can speed up developers and data scientists’ productivity. By designing data science experiments following a common framework and enabling push workflows without the need of switching to GitHub can save previous time, so that data scientists can focus on the core aspects of their projects.
## Marketing Assistant
As we saw in Chapter 6 – Mastering Marketing with ChatGPT, AI assistant can be extremely valuable when it comes to this domain. In fact, generating text content – like social media posts, blog articles or marketing campaigns – is probably the activity in which these models perform at their best.
> A copywriter is a professional who specializes in writing persuasive and engaging content, often for marketing and advertising purposes. Their work typically includes writing promotional materials such as advertisements, brochures, websites, emails, social media posts, and other forms of content designed to persuade an audience to take a specific action, like making a purchase or subscribing to a service.
In this section, we are going to create a Copywrite assistant which is tailored to this kind of activity. To do so, I named my assistant Copywrite Companion, and I set the following configuration components:
* Instructions:
*Copywrite Companion is a versatile assistant designed to help users with various writing tasks. It specializes in generating product sheets based on user input, including text and images, crafting email campaigns for newsletters or promotional purposes, creating engaging visuals for new products, and writing tailored social media posts for different platforms. It ensures content is persuasive, engaging, and aligned with the target audience, aiming to increase engagement and promote products or services effectively. The assistant adapts its style to the user's needs and strives for creativity, clarity, and relevance in all outputs. It communicates in a casual, friendly tone, making interactions feel approachable and easygoing.*
* Conversation starters:
*Can you write a product description for my new product?*
*I need a catchy headline for an article about climbing*
*Give me a few ideas for a blog post about running*
* Capabilities:

Figure 31: Copywrite Companion’s enabled plugins.
Let’s see it in action:
* I’ll first ask it to write a product sheet starting from a picture provided (output truncated):

Figure 32: Example of Copywrite Companion generating a product description.
* As copywriter, we might want to insert this set of information into more structured repository, like an excel file. Let’s ask the assistant to do so, leveraging its code interpreter plugin:

Figure 33: Example of Copywrite Companion converting its previous response into an excel file, leveraging the code interpreter plugin.
* And this is the final result:

Figure 34: Excel file generated by Copywrite Companion
Let’s now ask the assistant to generate an Instagram post to sponsor our shoes:

Figure 35: Example of Copywrite Companion generating an Instagram post.
As you can see, the assistant also proposed an image description to leverage the DALL-E plugin. Since the description makes sense to me, I’ll go ahead and ask it to generate it:

Figure 36 Example of Copywrite Companion generating an image leveraging the DALL-E plugin.
Finally, I want to have an idea of how big competitor brands – like Nike and Adidas – are doing their marketing activities. To do so, I’ll ask my assistant to gather some evidence from the web:

Figure 37: Example of Copywrite Companion doing a competitive analysis leveraging the web browsing plugin.
As you can see, our companion correctly leveraged the web browsing plugin to retrieve the required information. Plus, it gave us powerful insight into which leverages the two competitor companies are investing, so that we might think (or ask our companion) about unique differentiators that could let our brand outstand in a competitive market.
We could also do a step ahead and ask for more specific insights about the competition by leveraging the code interpreter and data analysis plugin. Let’s say, for example, that we gathered an excel sheet with the following structure:

Figure 38: Competitive analysis on an Excel sheet
And now we want to generate some visuals out of it. Let’s ask our Copywriter Companion to do so:

Figure 39: Example of bar chart and scatter plot generated by Copywrite companion.

Figure 40: Example of line chart generated by Copywrite companion.
Including, as requested, the executive report:

Figure 41: Example of an Executive report generated by Copywrite companion.
Overall, tailoring ChatGPT for marketing activities can be extremely useful when it comes to generating new content, designing marketing strategies and doing competitive analysis on the web.
## Research Assistant
In this scenario, we are going focus on research once more, yet this time with a particular focus on papers retrieval. More specifically, we want our assistant to be able to do the following:
* Retrieving information from custom knowledge base that we provide
* Integrating the custom documents with papers only from Arxiv, enabling the web plugin for this task
* Retrieving from a Database (in our case, it will be hosted in Notion) existing ongoing work from other researchers, so that we don’t risk elaborating an essay which is already covered by someone else.
Let’s see all the steps.
1. Upload custom documents. For this purpose, I’ll use two papers about Image Classification in Machine Learning: “*CIFAKE: image classification and explainable identification of ai-generated synthetic images*” from J. Bird et al., and “*A Comprehensive Study of Vision Transformers in Image Classification Tasks*” from Khalis et al.
You can upload them in the proper section in the configuration pane:

Figure 42: Uploading custom documents in the configuration pane
1. Integrating custom documents with web references. To do so, we need to enable the web plugin:

Figure 43: Enabling the web browsing plugin in the configuration pane
Plus, we also need to specify the assistant to only navigate through the Arxiv archive. We will see how to specify that in the set of instructions we will create.
1. Retrieve information from a Notion Database. Here, the idea is that, as researchers, we might come up with ideas that are already being studied and developed by other fellow colleagues. Imagine that we keep track of all the ongoing research in a Notion Database that has the following structure:

Figure 44: Notion Database structure.
To do so, we need to create a GPT actions. To do so, there are two steps to follow:
1. In your Notion workspace, you need to create a new connection marked as internal, along with a new Internal Integration Secret (or API key) will be created. You can call this connection “chatgpt” or similar.
2. In your GPT configuration pane, you need to set a new Action schema. Since in our case we need to query a specific Database, the schema will look like the following:
3. When it comes to authentication, you can click on "Authentication" and choose "API Key". Enter in the information below.
* API Key: Use Internal Integration Secret from your newly created connection in Notion
* Auth Type: Bearer

Figure 45: Notion Action Schema.
This is how the whole schema looks like:
openapi: 3.1.0
info:
title: Notion API
description: API for interacting with Notion's pages, databases, and users.
version: 1.0.0
servers:
- url: https://api.notion.com/v1
description: Main Notion API server
paths:
/databases/{database_id}/query:
post:
operationId: queryDatabase
summary: Query a database
parameters:
- name: database_id
in: path
required: true
schema:
type: string
- name: Notion-Version
in: header
required: true
schema:
type: string
example: 2022-06-28
constant: 2022-06-28
requestBody:
required: true
content:
application/json:
schema:
type: object
properties:
filter:
type: object
sorts:
type: array
items:
type: object
start_cursor:
type: string
page_size:
type: integer
responses:
'200':
description: Successful response
content:
application/json:
schema:
type: object
properties:
object:
type: string
results:
type: array
items:
type: object
next_cursor:
type: string
has_more:
type: boolean
Great, now that we have all our ingredients, we need to set our system message and, optionally, our conversation starters. In this scenario, I set the following instructions:
*You are an AI research assistant tasked with helping researchers by leveraging a variety of tools and resources. Your primary responsibilities include retrieving, integrating, and cross-referencing information from custom knowledge bases, academic papers on Arxiv, and ongoing research projects stored in a Notion database. Follow these guidelines to ensure your assistance is accurate, comprehensive, and avoids redundancy:*
1. *Always answer based on the provided documents*
2. *If you feel it is needed, extend your search on the web. the ONLY ONE SITE you can navigate through is Arxiv.*
3. *If asked by the user, check whether the topic is already covered in the Notion DB.*
And the following conversation starters:
1. *How do VGGNet, ResNet, and Inception differ in accuracy and efficiency for image classification?*
2. *What challenges do image classification models face, and how do techniques like data augmentation and transfer learning help?*
3. *How does CNN depth affect image classification, and how do ResNet's residual connections help?*
The final product looks like so:

Figure 46: Landing page of our ResearchGPT
Let’s test it:
1. I’ll first ask a generic question, that will be addressed with the provided papers:

Figure 47: Example of ResearchGPT retrieving knowledge from provided documents
2. Then, I want to see whether the topic is already covered:

Figure 48: Example of ResearchGPT talking to Notion with the pre-defined action.
3. Finally, I want to integrate the topic to make it unique:

Figure 49: Example of ResearchGPT using the web browser plugin.
As you can see, our Assistant could leverage all the tools we provided it with, calling the Notion DB when needed.
## Summary
In this Chapter, we explored how to get customized GPTs to address your specific goals. The possibility of tailoring ChatGPT opens a new landscape of scenarios, where highly specialized AI assistant becomes the everyday companion of professionals. Plus, OpenAI offered a no-code UI to create GPTs, so that all citizen developers can benefit not only from the extended marketplace of available solutions, but also from their own creations.
With the plugins extensibility and the custom knowledge base, custom GPTs can serve you for limitless tasks. Then, with the addition of powerful actions, they can also communicate with the surrounding environment and evolve from “mere” generation to automation.
With this Chapter, we also conclude Part 2 of this book, where we focused on practical applications of ChatGPT. Starting from next Chapter, we are going to cover more in details how large enterprises can leverage OpenAI models and embed them into their business processes. We will cover trending enterprise use cases and see practical implementation on how to embed OpenAI’s LLMs via REST API in your applications.
## References
* [https://arxiv.org/abs/2312.11914](https://arxiv.org/abs/2312.11914)
* [https://arxiv.org/abs/2101.07714](https://arxiv.org/abs/2101.07714)
* [https://arxiv.org/abs/2101.07714](https://arxiv.org/abs/2101.07714)
* [https://psycnet.apa.org/record/2014-07087-000](https://psycnet.apa.org/record/2014-07087-000)
* [https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109](https://www.tandfonline.com/doi/abs/10.2753/MIS0742-1222240109)
* [https://www.youtube.com/watch?v=TUmyygCMMGA](https://www.youtube.com/watch?v=TUmyygCMMGA)
* [https://arxiv.org/pdf/2312.01232](https://arxiv.org/pdf/2312.01232)
* [https://arxiv.org/pdf/2303.14126](https://arxiv.org/pdf/2303.14126)

















浙公网安备 33010602011771号