惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
V
Vulnerabilities – Threatpost
B
Blog
月光博客
月光博客
宝玉的分享
宝玉的分享
有赞技术团队
有赞技术团队
美团技术团队
IT之家
IT之家
B
Blog RSS Feed
V
V2EX
Hugging Face - Blog
Hugging Face - Blog
T
The Blog of Author Tim Ferriss
Vercel News
Vercel News
Jina AI
Jina AI
Y
Y Combinator Blog
Recorded Future
Recorded Future
N
Netflix TechBlog - Medium
S
SegmentFault 最新的问题
L
LangChain Blog
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
PCI Perspectives
PCI Perspectives
Schneier on Security
Schneier on Security
Microsoft Azure Blog
Microsoft Azure Blog
P
Privacy International News Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
C
Cyber Attacks, Cyber Crime and Cyber Security
N
News and Events Feed by Topic
W
WeLiveSecurity
L
Lohrmann on Cybersecurity
Security Archives - TechRepublic
Security Archives - TechRepublic
Help Net Security
Help Net Security
Google DeepMind News
Google DeepMind News
P
Proofpoint News Feed
S
Schneier on Security
Last Week in AI
Last Week in AI
L
LINUX DO - 最新话题
Webroot Blog
Webroot Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
云风的 BLOG
云风的 BLOG
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
H
Hackread – Cybersecurity News, Data Breaches, AI and More
C
CXSECURITY Database RSS Feed - CXSecurity.com
J
Java Code Geeks
T
Threatpost
腾讯CDC
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
The Exploit Database - CXSecurity.com
H
Help Net Security

博客园 - cleardo

大模型的原理学习(一) 我的测试开发十年之路 命令行安装ipa包 ios设备管理 tomcat远程部署 移动线路的测试方案设计 iOS开发过程中的内存监控 iOS添加图片 iOS交叉编译 iOS日志获取 使用flask开发web应用 Python隔离环境的搭建 Ipa打包并安装到iphone 单元测试的痛点 Graphviz入门 Iphone常用工具 iOS越狱后必装软件 构建iOS交叉编译环境 pycurl库使用详解
大模型的原理学习(二)
cleardo · 2026-02-04 · via 博客园 - cleardo

(一) 介绍

(二) 损失函数

(三) CNN

(四) RNN

(五) Transformer架构

(六) GPT

(七) 幻觉问题

(八) RAG

(九) MCP

(十) Skills

第一篇链接回顾: 大模型的原理学习(一)

在第二篇中,我们开始学习下损失函数和拟合的概念。

拟合与损失函数

如何判断拟合的好不好?

损失函数或者代价函数(损失函数的平均值)

拟合动画

损失函数:

1、 均方误差

均方误差

2、 二分类交叉熵

二分类交叉熵

如何优化?

梯度下降法

目标:使得代价函数尽可能小,那就要求偏导

函数 =
偏导

经过激活函数=
激活函数

① 计算损失函数 ,选用二分类交叉熵 得到

损失函数

② 对损失函数中的 w 和 b 求偏导,偏导的作用代表的就是参数变化最快的方向,设置一个步长(学习率 η)

步长

③ 更新参数:每经过一次计算,得到一个 W 和 b 就带入一下计算损失函数,看看损失函数有没有变小,如果变小了,说明本次 w 和 b 的调整是有效的,就更新 w 和 b

④ 继续重复 ①~③ 步骤,直到损失函数达到最小。

梯度下降

在神经网络中,是如何通过一层层隐藏函数,最终找出最优的输出?

前向传播:根据输入 x 计算出预测值 y

反向传播:计算预测值与实际值的损失函数,进行梯度优化

梯度下降2

链式法则

偏导链式法则

第三篇,我们将开始学习卷积神经网络的原理

卷积神经网络(Convolutional Neural Network, CNN)

前面我们都是以一个输入经过神经网络得到一个输出为例,那么当输入很多的时候,神经网络面临什么问题呢?