Pytorch（二） —— 激活函数、损失函数及其梯度

1.激活函数
2.损失函数
- 2.1 MSE
- 2.2 CorssEntorpy
3. 求导和反向传播
- 3.1 求导
- 3.2 反向传播

1.激活函数

1.1 Sigmoid / Logistic

$\delta(x)=\frac{1}{1+e^{-x}}\\\delta'(x)=\delta(1-\delta)$

import matplotlib.pyplot as plt
import torch.nn.functional as F
x = torch.linspace(-10,10,1000)
y = F.sigmoid(x)
plt.plot(x,y)
plt.show()
1
2
3
4
5
6

在这里插入图片描述

1.2 Tanh

$tanh(x)=\frac{e^x-e^{-x}}{e^x+e^{-x}}\\\frac{\partial tanh(x)}{\partial x}=1-tanh^2(x)$

import matplotlib.pyplot as plt
import torch.nn.functional as F
x = torch.linspace(-10,10,1000)
y = F.tanh(x)
plt.plot(x,y)
plt.show()
1
2
3
4
5
6

在这里插入图片描述

1.3 ReLU

$f (x) = m a x (0, x)$

import matplotlib.pyplot as plt
import torch.nn.functional as F
x = torch.linspace(-10,10,1000)
y = F.relu(x)
plt.plot(x,y)
plt.show()
1
2
3
4
5
6

在这里插入图片描述

1.4 Softmax

$p_i=\frac{e^{a_i}}{\sum_{k=1}^N{e^{a_k}}}\\ \frac{\partial p_i}{\partial a_j}=\left\{$

\begin{array}{lc} p_{i} (1 - p_{j}) & i = j \\ - p_{i} p_{j} & i \neq j \end{array}

\right.

p_{i} = \frac{e ^{a_{i}}}{\sum _{k = 1}^{N} e ^{a_{k}}} \frac{\partial p _{i}}{\partial a _{j}} = {p_{i} (1 - p_{j}) - p_{i} p_{j} i = j i \neq = j

import torch.nn.functional as F
logits = torch.rand(10)
prob = F.softmax(logits,dim=0)
print(prob)
1
2
3
4

tensor([0.1024, 0.0617, 0.1133, 0.1544, 0.1184, 0.0735, 0.0590, 0.1036, 0.0861,
        0.1275])
1
2

2.损失函数

2.1 MSE

import torch.nn.functional as F
x = torch.rand(100,64)
w = torch.rand(64,1)
y = torch.rand(100,1)
mse = F.mse_loss(y,x@w)
print(mse)
1
2
3
4
5
6

tensor(238.5115)
1

2.2 CorssEntorpy

import torch.nn.functional as F
x = torch.rand(100,64)
w = torch.rand(64,10)
y = torch.randint(0,9,[100])
entropy = F.cross_entropy(x@w,y)
print(entropy)
1
2
3
4
5
6

tensor(3.6413)
1

3. 求导和反向传播

3.1 求导

Tensor.requires_grad_()
torch.autograd.grad()

import torch.nn.functional as F
import torch
x = torch.rand(100,64)
w = torch.rand(64,1)
y = torch.rand(100,1)
w.requires_grad_()
mse = F.mse_loss(x@w,y)
grads = torch.autograd.grad(mse,[w])
print(grads[0].shape)
1
2
3
4
5
6
7
8
9

torch.Size([64, 1])
1

3.2 反向传播

Tensor.backward()

import torch.nn.functional as F
import torch
x = torch.rand(100,64)
w = torch.rand(64,10)
w.requires_grad_()
y = torch.randint(0,9,[100,])
entropy = F.cross_entropy(x@w,y)
entropy.backward()
w.grad.shape
1
2
3
4
5
6
7
8
9

torch.Size([64, 10])
1

by CyrusMay 2022 06 28

人生只是须臾的刹那
人间只是天地的夹缝
——————五月天（因为你所以我）——————

相关阅读:
kubernetes教程-基本学习环境配置
nRF52832 SDK15.3.0 基于ble_app_uart demo FreeRTOS移植
软件开发遵循的一些原则
CV每日论文--2024.6.25
【深度学习】ONNX模型快速部署【入门】
手动编译与安装Qt的子模块
vue3 provide inject
为什么我建议在复杂但是性能关键的表上所有查询都加上 force index
java标注
webpack内使用babel

原文地址：https://blog.csdn.net/Cyrus_May/article/details/125500584