心血来潮想手撕一下pytorch的源码,于是从手动编译pytorch开始。首先第一步,创建一个干净的环境吧:

1
conda create --name torch-test python=3.11

然后从github上拉取一下pytorch的源码,文件比较大,所以用一下国内的镜像源:

1
git clone https://gitee.com/mirrors/pytorch.git

进入pytorch目录,找一个网络好一点的地方,下载所需要的子模块的源码:

1
2
3
cd pytorch
git submodule sync
git submodule update --init --recursive

这一步及其容易断网出错,试了很多次,而且踩了一个坑:第一次失败的时候third_party目录下因为已经存在许多空目录,看着心烦所以我把third_party整个目录删掉了,但是不料third_party目录下有很多不是用过submodule命令拉取下来的文件和目录,导致我后续一直编译失败而又在.gitmodules文件中找不到报错对应的库。后来求助了万能的gemini大人,才发现删掉了很多不该删除的文件,然后从回收站里一个个的把他们捡回来QwQ。

这里简单分析一下pytorch所需的子模块,抽取部分来进行分析,对pytorch的构成有一个初步的印象。首先pybind11是一个用于提供python和c++之间的桥梁的库,由于pytorch的计算核心都是c++写的,而需要提供python接口,就需要pybind11来进行。googletest,benchmark等库都是用来进行c++代码的测试的,还有一些其他的库提供了c++开发的一些基础设施。

对于核心的计算部分,eigen是一个用于高效线性代数、矩阵计算的c++库,不过pytorch的核心张量计算并不依赖这个库。NNPACK是一个针对多核cpu高性能神经网络计算的库,还有许多其他的库是这个库的依赖项。sleef是一个高性能数学函数库,提供了许多数学函数的运算操作。这里在介绍一下flash-attention库,这是一种高效的注意力机制实现,目前许多大模型推理都有用到这个技术。

言归正传,下载好所有子模块后,在pytorch目录下安装相关依赖:

1
2
3
4
conda install cmake ninja
pip install -r requirements.txt
#为了使用torch.distributed,所以也安装一下!
conda install pkg-config libuv

简单分析一下所安装的依赖,其中cmake用于管理编译过程。cmake是一种跨平台的构建系统生成工具,其提供了一种高层次的抽象描述,使得开发人员不用关注每个平台具体的编译细节,而是通过编写CmakeLists.txt来描述编译的过程,配置,依赖等。而ninja是一种高效的构建工具,其语法设计非常简单,一般来说其命令不由人直接编写,而是由cmake等工具生成。其获取所有编译命令后,会对编译命令进行分析,然后进行并行编译(对于彼此之间没有依赖关系的编译命令)和增量编译(根据时间戳进行分析)。

安装完成后,就可以开始进行编译啦!运行:

1
python setup.py develop

等待大概50分钟,即可编译成功!如果中途因为什么缘故编译失败了,再次进行编译之前,最好清理一下上一次编译的中间产物,直接运行:

1
git clean -xdf

可以删除所有未被git所追踪的文件。

接下来,随便找个地方新建一个test.py测试一下所编译的pytorch是否能用吧!

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
# test.py
import torch
import torch.nn as nn
import torch.optim as optim

# 0. 检查 PyTorch 版本和设备
print(f"PyTorch Version: {torch.__version__}")

# 优先使用 MPS (Apple Silicon GPU),如果可用
if torch.backends.mps.is_available():
device = torch.device("mps")
print("Using MPS device (Apple Silicon GPU)")
elif torch.cuda.is_available(): # 虽然你是在 Mac 上,但保留这个以防万一或将来在其他地方运行
device = torch.device("cuda")
print("Using CUDA device")
else:
device = torch.device("cpu")
print("Using CPU device")

# 1. 准备数据 (XOR problem)
# 输入: [0,0], [0,1], [1,0], [1,1]
# 输出: 0, 1, 1, 0
X = torch.tensor([[0, 0],
[0, 1],
[1, 0],
[1, 1]], dtype=torch.float32).to(device)

Y = torch.tensor([[0],
[1],
[1],
[0]], dtype=torch.float32).to(device)

# 2. 定义神经网络模型
class SimpleNN(nn.Module):
def __init__(self, input_size, hidden_size, output_size):
super(SimpleNN, self).__init__()
self.fc1 = nn.Linear(input_size, hidden_size) # 输入层到隐藏层
self.relu = nn.ReLU() # 激活函数
self.fc2 = nn.Linear(hidden_size, output_size)# 隐藏层到输出层
self.sigmoid = nn.Sigmoid() # 输出层激活函数 (用于二分类问题)

def forward(self, x):
out = self.fc1(x)
out = self.relu(out)
out = self.fc2(out)
out = self.sigmoid(out) # 输出概率在 0 到 1 之间
return out

# 实例化模型
input_size = 2 # 输入特征数 (x1, x2)
hidden_size = 4 # 隐藏层神经元数量 (可以调整)
output_size = 1 # 输出神经元数量 (一个概率值)

model = SimpleNN(input_size, hidden_size, output_size).to(device)
print("\nModel Architecture:")
print(model)

# 3. 定义损失函数和优化器
criterion = nn.BCELoss() # Binary Cross Entropy Loss,适用于二分类问题
optimizer = optim.SGD(model.parameters(), lr=0.1) # 使用随机梯度下降优化器

print("\nInitial model parameters (first layer weights):")
print(model.fc1.weight)

# 4. 训练模型
epochs = 10000 # 训练轮数
print_every = 1000

print("\nStarting training...")
for epoch in range(epochs):
# 前向传播
outputs = model(X)
loss = criterion(outputs, Y)

# 反向传播和优化
optimizer.zero_grad() # 清空之前的梯度
loss.backward() # 计算梯度
optimizer.step() # 更新权重

if (epoch + 1) % print_every == 0:
print(f'Epoch [{epoch+1}/{epochs}], Loss: {loss.item():.4f}')

print("Training finished.")

print("\nFinal model parameters (first layer weights):")
print(model.fc1.weight)

# 5. 测试模型
with torch.no_grad(): # 在测试阶段不计算梯度
predicted = model(X)
# 将概率转换为 0 或 1 的预测
predicted_classes = (predicted > 0.5).float()
accuracy = (predicted_classes == Y).float().mean()
print(f'\nAccuracy on training data: {accuracy.item()*100:.2f}%')

print("\nPredictions vs True values:")
for i in range(len(X)):
input_val = X[i].cpu().numpy() # 转到 CPU 并转为 numpy 方便打印
true_val = Y[i].cpu().item()
pred_prob = predicted[i].cpu().item()
pred_class = predicted_classes[i].cpu().item()
print(f"Input: {input_val}, True: {true_val}, Predicted Prob: {pred_prob:.4f}, Predicted Class: {pred_class}")

print("\nTesting individual inputs:")
test_inputs = [
torch.tensor([0,0], dtype=torch.float32).to(device),
torch.tensor([0,1], dtype=torch.float32).to(device),
torch.tensor([1,0], dtype=torch.float32).to(device),
torch.tensor([1,1], dtype=torch.float32).to(device)
]

for test_input in test_inputs:
with torch.no_grad():
pred = model(test_input.unsqueeze(0)) # unsqueeze(0) 来增加 batch 维度
pred_class = (pred > 0.5).float().item()
print(f"Input: {test_input.cpu().numpy()}, Predicted Output: {pred_class}")

输出如下:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
PyTorch Version: 2.8.0a0+git3aa8477
Using MPS device (Apple Silicon GPU)

Model Architecture:
SimpleNN(
(fc1): Linear(in_features=2, out_features=4, bias=True)
(relu): ReLU()
(fc2): Linear(in_features=4, out_features=1, bias=True)
(sigmoid): Sigmoid()
)

Initial model parameters (first layer weights):
Parameter containing:
tensor([[ 0.3804, -0.3864],
[-0.4644, -0.3001],
[-0.1331, -0.4035],
[ 0.0199, 0.1429]], device='mps:0', requires_grad=True)

Starting training...
Epoch [1000/10000], Loss: 0.4809
Epoch [2000/10000], Loss: 0.4783
Epoch [3000/10000], Loss: 0.4779
Epoch [4000/10000], Loss: 0.4777
Epoch [5000/10000], Loss: 0.4776
Epoch [6000/10000], Loss: 0.4776
Epoch [7000/10000], Loss: 0.4775
Epoch [8000/10000], Loss: 0.4775
Epoch [9000/10000], Loss: 0.4775
Epoch [10000/10000], Loss: 0.4775
Training finished.

Final model parameters (first layer weights):
Parameter containing:
tensor([[ 0.3804, -0.3864],
[-2.5005, 2.4790],
[-0.1331, -0.4035],
[ 0.0199, 0.1429]], device='mps:0', requires_grad=True)

Accuracy on training data: 75.00%

Predictions vs True values:
Input: [0. 0.], True: 0.0, Predicted Prob: 0.3335, Predicted Class: 0.0
Input: [0. 1.], True: 1.0, Predicted Prob: 0.9996, Predicted Class: 1.0
Input: [1. 0.], True: 1.0, Predicted Prob: 0.3335, Predicted Class: 0.0
Input: [1. 1.], True: 0.0, Predicted Prob: 0.3335, Predicted Class: 0.0

Testing individual inputs:
Input: [0. 0.], Predicted Output: 0.0
Input: [0. 1.], Predicted Output: 1.0
Input: [1. 0.], Predicted Output: 0.0
Input: [1. 1.], Predicted Output: 0.0

完结撒花!