c85bbd3c55
- AGENTS.md: 新增 Codex 代理说明文件 - .gitignore: 屏蔽第三方工具目录 (.prism/, .claudeprism/, .agents/) - preamble: 启用 xcolor dvipsnames 选项,新增 GREEN 颜色与 \greenheading 命令 - ch02-framework: 修正智能体段落引号格式 - ap01-programming: 补充 Python 模块与包的对比说明(esp_engine 示例) - ap02-vibe-coding: 应用 \greenheading 标题样式,强调 commit 作为审核检查点 - 移除 2026_DesignAI.code-workspace(放弃 workspace 持久化方案) Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
1186 lines
35 KiB
TeX
1186 lines
35 KiB
TeX
\chapter{计算机基础与编程环境}
|
||
|
||
本附录介绍计算机的基本构成、操作系统常见操作、程序设计语言的概念、Python编程语言的系统学习指南,以及机器学习与深度学习的实践入门。
|
||
|
||
\subsection{计算机的基本构成与操作系统常见操作}
|
||
|
||
\subsubsection{计算机基本构成}
|
||
|
||
了解计算机的基本组成有助于理解AI程序运行时的资源需求。
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.2368}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.3026}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.4342}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
组件
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
作用
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
AI相关说明
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
CPU & 中央处理器,执行指令 & 数据预处理、逻辑控制 \\
|
||
GPU & 图形处理器,并行计算 & 深度学习训练与推理的核心硬件 \\
|
||
内存(RAM) & 临时存储运行中的数据 & 影响能处理的批量大小和数据规模 \\
|
||
硬盘(SSD/HDD) & 持久化存储数据 & 模型文件、数据集的存储 \\
|
||
网络 & 数据传输 & 下载模型、调用云端API \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{操作系统常见操作}
|
||
|
||
本书以 macOS/Linux 为主要环境,Windows 用户推荐使用 WSL2(Windows
|
||
Subsystem for Linux)。
|
||
|
||
\paragraph{文件与目录操作:}
|
||
|
||
\emph{\# 查看当前路径}\\
|
||
pwd\\
|
||
\strut \\
|
||
\emph{\# 列出文件}\\
|
||
ls -la\\
|
||
\strut \\
|
||
\emph{\# 创建目录}\\
|
||
mkdir my\_project\\
|
||
\strut \\
|
||
\emph{\# 切换目录}\\
|
||
cd my\_project\\
|
||
\strut \\
|
||
\emph{\# 复制、移动、删除}\\
|
||
cp file.txt backup.txt\\
|
||
mv old.txt new.txt\\
|
||
rm unwanted.txt
|
||
|
||
\paragraph{环境与进程管理:}
|
||
|
||
\emph{\# 查看系统资源}\\
|
||
top \emph{\# CPU和内存使用}\\
|
||
df -h \emph{\# 磁盘空间}\\
|
||
nvidia-smi \emph{\# GPU状态(NVIDIA显卡)}\\
|
||
\strut \\
|
||
\emph{\# 包管理}\\
|
||
brew install xxx \emph{\# macOS Homebrew}\\
|
||
apt install xxx \emph{\# Ubuntu/Debian}
|
||
|
||
\subsubsection{硬件资源推荐}
|
||
|
||
\textbf{本地GPU配置}: - GPU:RTX 3060 (12GB) 或更高 - 内存:16GB+ -
|
||
存储:至少100GB SSD
|
||
|
||
\textbf{云平台}:
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.2639}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1528}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1528}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
平台
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
特点
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
适合场景
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
Google Colab & 免费GPU & 学习实验 \\
|
||
Kaggle Notebooks & 免费GPU & 竞赛 \\
|
||
AutoDL & 按时计费 & 中期项目 \\
|
||
阿里云PAI & 国内稳定 & 生产部署 \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{软件安装与运行}
|
||
|
||
\textbf{``安装''的本质就是}:把编译好的二进制文件放到 PATH
|
||
某个目录下,让 shell 能找到它。例如git的安装和使用:
|
||
\textbf{总结:三层抽象}
|
||
|
||
┌─────────────────────────────────────────────┐\\
|
||
│ 用户层:brew install git / git clone │ ← 你看到的\\
|
||
├─────────────────────────────────────────────┤\\
|
||
│ Shell 层:搜索 PATH → execve() 加载二进制 │ ← 为什么能找到命令\\
|
||
├─────────────────────────────────────────────┤\\
|
||
│ OS 层:系统调用 (open/read/write/socket) │ ← 为什么能真正干活\\
|
||
├─────────────────────────────────────────────┤\\
|
||
│ 硬件层:CPU 执行指令、网卡收发数据、磁盘写入 │ ← 物理上发生了什么\\
|
||
└─────────────────────────────────────────────┘
|
||
|
||
所以整个链条是:\textbf{包管理器下载编译好的二进制 → 放到 PATH 目录 →
|
||
shell 通过 PATH 找到它 → execve 加载到内存 →}
|
||
|
||
\textbf{二进制内部调用 OS API
|
||
完成实际工作}。没有任何“魔法”,本质上就是文件操作和进程管理的组合。
|
||
|
||
\subsection{程序设计语言与软件开发}
|
||
|
||
\subsubsection{什么是程序设计语言}
|
||
|
||
计算机只能执行由0和1组成的\textbf{机器码}(machine
|
||
code),但人类直接阅读和编写机器码极其困难。程序设计语言就是人与计算机之间的桥梁——用人类可读的语法表达逻辑,再通过特定工具转换为机器可执行的指令。
|
||
|
||
从底层到高层的演进:
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1429}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.3506}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.4805}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
层级
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
语言示例
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
特点
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
机器码 & 二进制 01101000 & 计算机直接执行,人类不可读 \\
|
||
汇编语言 & MOV AX, 1 & 与机器码一一对应,可读性低 \\
|
||
低级语言 & C & 接近硬件,性能高,需要手动管理内存 \\
|
||
高级语言 & Python, Java, JavaScript & 接近自然语言,开发效率高 \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{编译型与解释型语言}
|
||
|
||
高级语言需要转换为机器码才能运行,根据转换方式的不同,分为两大类:
|
||
|
||
\textbf{编译型语言(Compiled)}:程序编写完成后,通过编译器一次性将全部代码翻译成机器码,生成可执行文件。
|
||
|
||
源代码 → 编译器 → 可执行文件 → 运行
|
||
|
||
C/C++:系统级开发、高性能计算、游戏引擎
|
||
|
||
Go:云服务、容器工具(Docker 即用 Go 编写)
|
||
|
||
Rust:系统编程,兼顾性能与安全
|
||
|
||
特点:运行速度快,但每次修改代码都需要重新编译。
|
||
|
||
\textbf{解释型语言(Interpreted)}:程序运行时,由解释器逐行读取代码并即时执行,不需要预先编译。
|
||
|
||
源代码 → 解释器逐行执行
|
||
|
||
Python:AI/数据科学的首选语言
|
||
|
||
JavaScript:网页交互、前端开发
|
||
|
||
Ruby:Web开发(Ruby on Rails)
|
||
|
||
特点:开发灵活、调试方便,但运行速度通常慢于编译型语言。
|
||
|
||
\textbf{混合模式}:Java 采用“编译为字节码 →
|
||
虚拟机解释执行”的混合方式,兼顾了跨平台和性能。
|
||
|
||
\subsubsection{常见编程语言概览}
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1512}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1047}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.3605}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.3605}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
语言
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
类型
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
主要用途
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
与AI/设计的关系
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
Python & 解释型 & AI、数据科学、自动化 & 本书主要编程语言 \\
|
||
C/C++ & 编译型 & 操作系统、嵌入式、高性能计算 &
|
||
深度学习框架的底层实现 \\
|
||
Java & 混合型 & 企业应用、Android开发 & 大数据处理(Hadoop/Spark) \\
|
||
JavaScript & 解释型 & 网页前端、Node.js后端 & Web可视化、交互设计 \\
|
||
Shell/Bash & 解释型 & 命令行脚本、系统管理 & 自动化任务、环境管理 \\
|
||
SQL & 声明式 & 数据库查询 & 数据管理与提取 \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\textbf{命令行界面(CLI)}:CLI(Command Line
|
||
Interface)是通过文本命令与计算机交互的方式。终端中输入的每一条命令(如
|
||
ls、git commit)本质上都是调用某个程序。掌握 CLI
|
||
是进行AI开发的基础技能,许多工具(如
|
||
conda、pip、git)主要通过命令行操作。
|
||
|
||
\subsubsection{软件开发基本概念}
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1327}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.3186}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.5310}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
概念
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
英文
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
说明
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
集成开发环境 & IDE & 集成代码编辑、调试、运行的开发工具(如 VS
|
||
Code、PyCharm) \\
|
||
编译器 & Compiler & 将源代码翻译为机器码的程序 \\
|
||
解释器 & Interpreter & 逐行读取并执行源代码的程序 \\
|
||
调试器 & Debugger & 帮助定位和修复代码错误的工具 \\
|
||
包管理器 & Package Manager & 管理第三方库的安装和更新(如
|
||
pip、conda、npm) \\
|
||
API & Application Programming Interface &
|
||
程序之间交互的接口(如调用AI模型的API) \\
|
||
开源 & Open Source & 源代码公开,可自由使用和修改 \\
|
||
版本控制 & Version Control & 管理代码的修改历史(如 Git) \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{Python环境配置}
|
||
|
||
Python是本书使用的核心编程语言,以下介绍环境搭建方法。
|
||
|
||
Anaconda vs Miniconda
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1165}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.0971}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1262}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.6408}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
工具
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
大小
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
特点
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
下载地址
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
Anaconda & \textasciitilde500MB & 预装常用库 &
|
||
\href{https://www.anaconda.com/download}{anaconda.com} \\
|
||
Miniconda & \textasciitilde50MB & 精简安装 &
|
||
\href{https://docs.conda.io/en/latest/miniconda.html}{docs.conda.io} \\
|
||
\end{longtable}
|
||
}
|
||
|
||
安装步骤
|
||
|
||
\# 1. 下载并安装 Anaconda 或 Miniconda\\
|
||
\strut \\
|
||
\# 2. 创建虚拟环境\\
|
||
conda create -n ai-env python=3.10\\
|
||
conda activate ai-env\\
|
||
\strut \\
|
||
\# 3. 安装核心科学计算库\\
|
||
conda install numpy pandas scipy\\
|
||
\strut \\
|
||
\# 4. 安装深度学习框架\\
|
||
pip install torch torchvision
|
||
|
||
\subsection{Python编程语言}
|
||
|
||
\paragraph{Python概述与特点}
|
||
|
||
Python由Guido van
|
||
Rossum于1991年发布,以“优雅”\,``简洁“``可读性强”为设计哲学。它是当前AI和数据科学领域使用最广泛的编程语言。
|
||
|
||
\textbf{核心特点}: - 语法简洁,接近自然语言,入门门槛低 -
|
||
丰富的第三方库生态(AI、数据处理、Web等) -
|
||
跨平台运行(Windows、macOS、Linux) - 活跃的开源社区支持
|
||
|
||
\paragraph{基础语法}
|
||
|
||
变量与数据类型
|
||
|
||
\emph{\# 变量赋值(无需声明类型)}\\
|
||
name = "设计人工智能" \emph{\# 字符串 str}\\
|
||
version = 1.0 \emph{\# 浮点数 float}\\
|
||
chapter\_count = 26 \emph{\# 整数 int}\\
|
||
is\_published = True \emph{\# 布尔值 bool}\\
|
||
\strut \\
|
||
\emph{\# 查看类型}\\
|
||
print(type(name)) \emph{\# \textless class
|
||
\textquotesingle str\textquotesingle\textgreater{}}
|
||
|
||
字符串操作
|
||
|
||
title = "设计人工智能"\\
|
||
\strut \\
|
||
\emph{\# 字符串拼接}\\
|
||
full\_title = title + ":基础与应用"\\
|
||
\strut \\
|
||
\emph{\# 格式化输出}\\
|
||
print(f"本书名为《\{full\_title\}》,共\{chapter\_count\}章")\\
|
||
\strut \\
|
||
\emph{\# 常用方法}\\
|
||
print(title.lower()) \emph{\# 设小写}\\
|
||
print(title.replace("人工智能", "AI")) \emph{\# 替换}\\
|
||
print(len(title)) \emph{\# 长度}
|
||
|
||
注释
|
||
|
||
\emph{\# 这是单行注释}\\
|
||
\strut \\
|
||
\emph{"""}\\
|
||
\emph{这是多行注释(文档字符串)}\\
|
||
\emph{常用于函数和类的说明}\\
|
||
\emph{"""}\\
|
||
\strut \\
|
||
\emph{\# 好的注释解释"为什么",而不是"做什么"}
|
||
|
||
\paragraph{数据结构}
|
||
|
||
列表(List)
|
||
|
||
有序、可变的序列,最常用的数据结构。
|
||
|
||
\emph{\# 创建列表}\\
|
||
models = {[}"CNN", "RNN", "Transformer", "Diffusion"{]}\\
|
||
\strut \\
|
||
\emph{\# 访问元素(索引从0开始)}\\
|
||
print(models{[}0{]}) \emph{\# CNN}\\
|
||
print(models{[}-1{]}) \emph{\# Diffusion(倒数第一个)}\\
|
||
\strut \\
|
||
\emph{\# 修改}\\
|
||
models.append("GAN") \emph{\# 添加元素}\\
|
||
models.remove("RNN") \emph{\# 删除元素}\\
|
||
models.sort() \emph{\# 排序}\\
|
||
\strut \\
|
||
\emph{\# 切片}\\
|
||
print(models{[}1:3{]}) \emph{\# 第2到第3个元素}
|
||
|
||
字典(Dictionary)
|
||
|
||
键值对结构,用于存储映射关系。
|
||
|
||
\emph{\# 创建字典}\\
|
||
model\_info = \{\\
|
||
"name": "ResNet",\\
|
||
"year": 2015,\\
|
||
"layers": 152,\\
|
||
"task": "图像分类"\\
|
||
\}\\
|
||
\strut \\
|
||
\emph{\# 访问}\\
|
||
print(model\_info{[}"name"{]}) \emph{\# ResNet}\\
|
||
\strut \\
|
||
\emph{\# 添加/修改}\\
|
||
model\_info{[}"accuracy"{]} = 0.96\\
|
||
\strut \\
|
||
\emph{\# 遍历}\\
|
||
\textbf{for} key, value \textbf{in} model\_info.items():\\
|
||
print(f"\{key\}: \{value\}")
|
||
|
||
元组(Tuple)
|
||
|
||
有序、不可变的序列,适合存储不变的数据。
|
||
|
||
\emph{\# 创建元组}
|
||
|
||
rgb = (255, 128, 0)
|
||
|
||
\emph{\# 解包}
|
||
|
||
\begin{enumerate}
|
||
\def\labelenumi{\arabic{enumi}.}
|
||
\tightlist
|
||
\item
|
||
g, b = rgb
|
||
\end{enumerate}
|
||
|
||
集合(Set)
|
||
|
||
无序、不重复的元素集合。
|
||
|
||
tools\_a = \{"Photoshop", "Figma", "Sketch"\}\\
|
||
tools\_b = \{"Figma", "Blender", "Rhino"\}\\
|
||
\strut \\
|
||
\emph{\# 集合运算}\\
|
||
print(tools\_a \& tools\_b) \emph{\# 交集: \{"Figma"\}}\\
|
||
print(tools\_a \textbar{} tools\_b) \emph{\# 并集}
|
||
|
||
\paragraph{控制流}
|
||
|
||
条件判断
|
||
|
||
loss = 0.05\\
|
||
\strut \\
|
||
\textbf{if} loss \textless{} 0.01:\\
|
||
print("模型收敛良好")\\
|
||
\textbf{elif} loss \textless{} 0.1:\\
|
||
print("模型基本收敛")\\
|
||
\textbf{else}:\\
|
||
print("模型需要继续训练")
|
||
|
||
循环
|
||
|
||
\emph{\# for 循环}\\
|
||
epochs = {[}1, 2, 3, 4, 5{]}\\
|
||
\textbf{for} epoch \textbf{in} epochs:\\
|
||
print(f"训练第 \{epoch\} 轮")\\
|
||
\strut \\
|
||
\emph{\# range 生成序列}\\
|
||
\textbf{for} i \textbf{in} range(10):\\
|
||
print(f"第\{i\}次迭代")\\
|
||
\strut \\
|
||
\emph{\# while 循环}\\
|
||
loss = 1.0\\
|
||
\textbf{while} loss \textgreater{} 0.01:\\
|
||
loss = loss * 0.9 \emph{\# 模拟训练过程}
|
||
|
||
列表推导式
|
||
|
||
简洁地创建新列表:
|
||
|
||
\emph{\# 传统写法}\\
|
||
squares = {[}{]}\\
|
||
\textbf{for} x \textbf{in} range(10):\\
|
||
squares.append(x ** 2)\\
|
||
\strut \\
|
||
\emph{\# 列表推导式(推荐)}\\
|
||
squares = {[}x ** 2 \textbf{for} x \textbf{in} range(10){]}\\
|
||
\strut \\
|
||
\emph{\# 带条件过滤}\\
|
||
even\_squares = {[}x ** 2 \textbf{for} x \textbf{in} range(10)
|
||
\textbf{if} x \% 2 == 0{]}
|
||
|
||
\paragraph{函数与模块}
|
||
|
||
定义函数
|
||
|
||
\textbf{def} calculate\_accuracy(correct, total):\\
|
||
\emph{"""计算准确率"""}\\
|
||
\textbf{return} correct / total\\
|
||
\strut \\
|
||
\emph{\# 调用函数}\\
|
||
acc = calculate\_accuracy(95, 100)\\
|
||
print(f"准确率: \{acc:.2\%\}")
|
||
|
||
默认参数与关键字参数
|
||
|
||
\textbf{def} train\_model(epochs, lr=0.001, optimizer="Adam"):\\
|
||
print(f"训练 \{epochs\} 轮, 学习率=\{lr\}, 优化器=\{optimizer\}")\\
|
||
\strut \\
|
||
\emph{\# 多种调用方式}\\
|
||
train\_model(100)\\
|
||
train\_model(100, lr=0.01)\\
|
||
train\_model(100, optimizer="SGD", lr=0.1)
|
||
|
||
导入模块
|
||
|
||
Python 通过\textbf{模块}(module)和\textbf{包}(package)组织代码复用:一个 \texttt{.py} 文件就是一个模块,含有 \texttt{\_\_init\_\_.py} 的文件夹就是一个包,使用 \texttt{from 包.包.模块 import 函数/类} 的语法导入。
|
||
|
||
以本书生态安全格局项目中的 \texttt{esp\_engine/} 为例:
|
||
|
||
\begin{verbatim}
|
||
esp_engine/
|
||
├── data_manager.py ← 模块(1个文件,7个函数)
|
||
├── io/ ← 包(文件夹,2个模块)
|
||
│ ├── __init__.py
|
||
│ ├── raster_io.py ← 7个函数
|
||
│ └── vector_io.py ← 5个函数
|
||
├── steps/ ← 包(文件夹,10个模块)
|
||
│ ├── __init__.py
|
||
│ ├── reclassify.py
|
||
│ ├── cost_distance.py
|
||
│ └── ...(8个文件)
|
||
└── workflows/ ← 包(文件夹,6个模块)
|
||
├── __init__.py
|
||
├── suitability.py
|
||
└── ...(5个文件)
|
||
\end{verbatim}
|
||
|
||
何时用模块、何时用包?原则很简单——看体量(表\ref{tab:module-vs-package})。
|
||
|
||
\begin{table}[htbp]
|
||
\centering
|
||
\caption{模块与包的选择原则}
|
||
\label{tab:module-vs-package}
|
||
\begin{tabular}{@{} l l l @{}}
|
||
\toprule
|
||
\textbf{情况} & \textbf{做法} & \textbf{项目中的例子} \\
|
||
\midrule
|
||
功能集中,一个文件装得下 & 模块(单个 .py) & data\_manager.py:7个函数全是数据预处理 \\[3pt]
|
||
功能分多个领域,各有独立逻辑 & 包(文件夹+子模块) & steps/:10个文件,每个是一种 GIS 操作 \\[3pt]
|
||
拿捏不准 & 先做模块,膨胀了再拆 & — \\
|
||
\bottomrule
|
||
\end{tabular}
|
||
\end{table}
|
||
|
||
一句话:\textbf{模块是``够用就好的最小单位'',包是``需要分家时的组织方式''}。
|
||
|
||
\emph{\# 导入标准库}
|
||
|
||
\textbf{import} os
|
||
|
||
\textbf{import} json
|
||
|
||
\emph{\# 导入第三方库}
|
||
|
||
\textbf{import} numpy \textbf{as} np
|
||
|
||
\textbf{import} pandas \textbf{as} pd
|
||
|
||
\emph{\# 从模块中导入特定功能}
|
||
|
||
\textbf{from} pathlib \textbf{import} Path
|
||
|
||
\textbf{from} collections \textbf{import} Counter
|
||
|
||
\emph{\# 安装第三方库(在终端中执行)}
|
||
|
||
\begin{lstlisting}
|
||
# pip install numpy pandas matplotlib
|
||
\end{lstlisting}
|
||
|
||
\paragraph{3.6 文件读写}
|
||
|
||
\emph{\# 读取文件}
|
||
|
||
\textbf{with} open("data.txt", "r", encoding="utf-8") \textbf{as} f:
|
||
|
||
content = f.read()
|
||
|
||
\emph{\# 逐行读取}
|
||
|
||
\textbf{with} open("data.csv", "r", encoding="utf-8") \textbf{as} f:
|
||
|
||
\textbf{for} line \textbf{in} f:
|
||
|
||
print(line.strip())
|
||
|
||
\emph{\# 写入文件}
|
||
|
||
\textbf{with} open("output.txt", "w", encoding="utf-8") \textbf{as} f:
|
||
|
||
f.write("分析结果\textbackslash n")
|
||
|
||
f.write(f"准确率: \{acc:.4f\}\textbackslash n")
|
||
|
||
\emph{\# 读写JSON(AI应用中常用的数据格式)}
|
||
|
||
\textbf{import} json
|
||
|
||
data = \{"model": "ResNet", "accuracy": 0.96\}
|
||
|
||
\emph{\# 写入JSON}
|
||
|
||
\textbf{with} open("result.json", "w") \textbf{as} f:
|
||
|
||
json.dump(data, f, indent=2)
|
||
|
||
\emph{\# 读取JSON}
|
||
|
||
\textbf{with} open("result.json", "r") \textbf{as} f:
|
||
|
||
\begin{lstlisting}
|
||
loaded = json.load(f)
|
||
\end{lstlisting}
|
||
|
||
\paragraph{面向对象编程基础}
|
||
|
||
面向对象编程(OOP)是Python的重要范式,许多AI库都基于OOP设计。
|
||
|
||
\textbf{class} NeuralNetwork:\\
|
||
\emph{"""简单的神经网络类"""}\\
|
||
\strut \\
|
||
\textbf{def} \_\_init\_\_(self, input\_size, hidden\_size,
|
||
output\_size):\\
|
||
\emph{"""初始化网络结构"""}\\
|
||
self.input\_size = input\_size\\
|
||
self.hidden\_size = hidden\_size\\
|
||
self.output\_size = output\_size\\
|
||
self.loss\_history = {[}{]}\\
|
||
\strut \\
|
||
\textbf{def} forward(self, x):\\
|
||
\emph{"""前向传播"""}\\
|
||
\emph{\# 这里简化为概念演示}\\
|
||
\textbf{return} f"输出:
|
||
输入\{x\}经过\{self.hidden\_size\}个隐藏层神经元"\\
|
||
\strut \\
|
||
\textbf{def} train(self, data, epochs=10):\\
|
||
\emph{"""训练网络"""}\\
|
||
\textbf{for} epoch \textbf{in} range(epochs):\\
|
||
loss = 1.0 / (epoch + 1) \emph{\# 模拟损失下降}\\
|
||
self.loss\_history.append(loss)\\
|
||
print(f"Epoch \{epoch+1\}, Loss: \{loss:.4f\}")\\
|
||
\strut \\
|
||
\emph{\# 创建实例}\\
|
||
model = NeuralNetwork(input\_size=784, hidden\_size=128,
|
||
output\_size=10)\\
|
||
\strut \\
|
||
\emph{\# 使用}\\
|
||
output = model.forward({[}0.5, 0.3, 0.8{]})\\
|
||
model.train(data=None, epochs=5)
|
||
|
||
\subsubsection{核心科学计算库}
|
||
|
||
NumPy:数值计算基础
|
||
|
||
\textbf{import} numpy \textbf{as} np\\
|
||
\strut \\
|
||
\emph{\# 创建数组}\\
|
||
x = np.array({[}1, 2, 3, 4{]})\\
|
||
\strut \\
|
||
\emph{\# 矩阵运算}\\
|
||
W = np.random.randn(4, 3) \emph{\# 4×3 权重矩阵}\\
|
||
h = np.dot(x, W) \emph{\# 矩阵乘法}\\
|
||
\strut \\
|
||
\emph{\# 激活函数}\\
|
||
relu = np.maximum(0, h) \emph{\# ReLU}\\
|
||
\strut \\
|
||
\emph{\# 统计运算}\\
|
||
print(np.mean(x)) \emph{\# 均值}\\
|
||
print(np.std(x)) \emph{\# 标准差}\\
|
||
print(np.max(x)) \emph{\# 最大值}
|
||
|
||
\paragraph{Pandas:数据处理}
|
||
|
||
\textbf{import} pandas \textbf{as} pd\\
|
||
\strut \\
|
||
\emph{\# 读取数据}\\
|
||
df = pd.read\_csv(\textquotesingle data.csv\textquotesingle)\\
|
||
\strut \\
|
||
\emph{\# 查看数据}\\
|
||
print(df.head()) \emph{\# 前几行}\\
|
||
print(df.shape) \emph{\# 行数和列数}\\
|
||
print(df.columns) \emph{\# 列名}\\
|
||
\strut \\
|
||
\emph{\# 数据清洗}\\
|
||
df = df.dropna() \emph{\# 删除缺失值}\\
|
||
\strut \\
|
||
\emph{\# 统计分析}\\
|
||
print(df.describe()) \emph{\# 描述性统计}\\
|
||
\strut \\
|
||
\emph{\# 筛选数据}\\
|
||
filtered = df{[}df{[}\textquotesingle accuracy\textquotesingle{]}
|
||
\textgreater{} 0.9{]}
|
||
|
||
\paragraph{Matplotlib:数据可视化}
|
||
|
||
\textbf{import} matplotlib.pyplot \textbf{as} plt\\
|
||
\strut \\
|
||
\emph{\# 折线图:训练损失曲线}\\
|
||
epochs = {[}1, 2, 3, 4, 5{]}\\
|
||
losses = {[}0.8, 0.5, 0.3, 0.15, 0.08{]}\\
|
||
\strut \\
|
||
plt.plot(epochs, losses, \textquotesingle b-o\textquotesingle)\\
|
||
plt.xlabel(\textquotesingle Epoch\textquotesingle)\\
|
||
plt.ylabel(\textquotesingle Loss\textquotesingle)\\
|
||
plt.title(\textquotesingle Training Loss Curve\textquotesingle)\\
|
||
plt.savefig(\textquotesingle loss\_curve.png\textquotesingle, dpi=150)\\
|
||
plt.show()
|
||
|
||
\subsection{AI工具库速查与模型资源}
|
||
|
||
\subsubsection{工具库速查}
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1340}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.1443}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.4227}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 6\tabcolsep) * \real{0.2784}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
类别
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
工具
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
安装命令
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
功能
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
计算机视觉 & OpenCV & pip install opencv-python & 图像处理、视频分析 \\
|
||
目标检测 & Ultralytics & pip install ultralytics &
|
||
YOLO目标检测、实例分割 \\
|
||
图像生成 & Diffusers & pip install diffusers & Stable
|
||
Diffusion模型调用 \\
|
||
Agent开发 & LangChain & pip install langchain langchain-openai &
|
||
LLM应用开发框架 \\
|
||
Agent开发 & LangGraph & pip install langgraph & 状态机式Agent开发 \\
|
||
数据检索 & LlamaIndex & pip install llama-index &
|
||
数据索引与检索(RAG) \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{模型资源}
|
||
|
||
\paragraph{\texorpdfstring{Hugging
|
||
Face:\href{https://huggingface.co/}{huggingface.co}}{Hugging Face:huggingface.co}}
|
||
|
||
功能:模型仓库、数据集、Spaces在线演示
|
||
|
||
\textbf{常用预训练模型}:
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1625}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.2750}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.5375}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
任务
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
推荐模型
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
Hugging Face ID
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
文生图 & Stable Diffusion XL &
|
||
stabilityai/stable-diffusion-xl-base-1.0 \\
|
||
目标检测 & YOLOv8 & Ultralytics \\
|
||
语义分割 & SAM & segment-anything \\
|
||
大语言模型 & Llama 3 & meta-llama/Meta-Llama-3-8B \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\section{机器学习与深度学习入门教程}
|
||
|
||
本节通过一个完整的实践流程,带领读者从数据准备到模型训练,体验机器学习和深度学习的核心步骤。
|
||
|
||
\subsubsection{机器学习 vs 深度学习}
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1358}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.4198}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.4198}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
维度
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
机器学习
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
深度学习
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
代表算法 & 线性回归、决策树、SVM、随机森林 & CNN、RNN、Transformer \\
|
||
特征工程 & 需要人工设计和选择特征 & 自动从原始数据中学习特征 \\
|
||
数据需求 & 中小规模数据即可 & 通常需要大量数据 \\
|
||
计算资源 & CPU即可 & 通常需要GPU \\
|
||
适用场景 & 结构化数据分析、基线模型 & 图像、文本、语音等非结构化数据 \\
|
||
\end{longtable}
|
||
}
|
||
|
||
\subsubsection{scikit-learn:机器学习实践}
|
||
|
||
scikit-learn是Python最经典的机器学习库,提供了丰富的算法和工具。
|
||
|
||
\paragraph{安装}
|
||
|
||
pip install scikit-learn
|
||
|
||
\paragraph{完整示例:鸢尾花分类}
|
||
|
||
\textbf{import} numpy \textbf{as} np
|
||
|
||
\textbf{import} matplotlib.pyplot \textbf{as} plt
|
||
|
||
\textbf{from} sklearn \textbf{import} datasets
|
||
|
||
\textbf{from} sklearn.model\_selection \textbf{import}
|
||
train\_test\_split
|
||
|
||
\textbf{from} sklearn.preprocessing \textbf{import} StandardScaler
|
||
|
||
\textbf{from} sklearn.linear\_model \textbf{import} LogisticRegression
|
||
|
||
\textbf{from} sklearn.tree \textbf{import} DecisionTreeClassifier
|
||
|
||
\textbf{from} sklearn.ensemble \textbf{import} RandomForestClassifier
|
||
|
||
\textbf{from} sklearn.metrics \textbf{import} accuracy\_score,
|
||
classification\_report
|
||
|
||
|
||
|
||
\emph{\# 1. 加载数据}
|
||
|
||
iris = datasets.load\_iris()
|
||
|
||
X = iris.data \emph{\# 特征:花萼长度、宽度,花瓣长度、宽度}
|
||
|
||
y = iris.target \emph{\# 标签:三种鸢尾花}
|
||
|
||
print(f"数据集大小: \{X.shape\}, 类别数: \{len(np.unique(y))\}")
|
||
|
||
\emph{\# 2. 划分训练集和测试集}
|
||
|
||
X\_train, X\_test, y\_train, y\_test = train\_test\_split(
|
||
|
||
\begin{lstlisting}
|
||
X, y, test_size=0.3, random_state=42
|
||
\end{lstlisting}
|
||
|
||
)
|
||
|
||
\emph{\# 3. 数据标准化}
|
||
|
||
scaler = StandardScaler()
|
||
|
||
X\_train = scaler.fit\_transform(X\_train)
|
||
|
||
X\_test = scaler.transform(X\_test)
|
||
|
||
\emph{\# 4. 训练多个模型并比较}
|
||
|
||
models = \{
|
||
|
||
"逻辑回归": LogisticRegression(),
|
||
|
||
\begin{quote}
|
||
"决策树": DecisionTreeClassifier(max\_depth=3),
|
||
|
||
"随机森林": RandomForestClassifier(n\_estimaors=100),
|
||
\end{quote}
|
||
|
||
\begin{lstlisting}
|
||
}
|
||
\end{lstlisting}
|
||
|
||
\textbf{for} name, model \textbf{in} models.items():
|
||
|
||
model.it(X\_train, y\_train)
|
||
|
||
y\_pred = model.predict(X\_test)
|
||
|
||
acc = accuracy\_score(y\_test, y\_pred)
|
||
|
||
print(f"\{name\} 准确率: \{acc:.2\%\}")
|
||
|
||
\emph{\# 5. 详细评估报告(以随机森林为例)}
|
||
|
||
best\_model = models{[}"随机森林"{]}
|
||
|
||
y\_pred = best\_model.predict(X\_test)
|
||
|
||
print("\textbackslash n分类报告:")
|
||
|
||
\begin{lstlisting}
|
||
print(classification_report(y_test, y_pred, target_names=iris.target_names))
|
||
\end{lstlisting}
|
||
|
||
\paragraph{机器学习工作流总结}
|
||
|
||
数据收集 → 数据预处理 → 特征工程 → 划分训练/测试集\\
|
||
→ 选择模型 → 训练 → 评估 → 调优 → 部署
|
||
|
||
\subsubsection{PyTorch:深度学习实践}
|
||
|
||
PyTorch是当前研究和实验中最流行的深度学习框架,以动态计算图和Pythonic
|
||
API著称。
|
||
|
||
\paragraph{张量(Tensor)基础}
|
||
|
||
张量是PyTorch的核心数据结构,类似于NumPy数组,但可以在GPU上运算。
|
||
|
||
\textbf{import} torch\\
|
||
\strut \\
|
||
\emph{\# 创建张量}\\
|
||
a = torch.tensor({[}1.0, 2.0, 3.0{]})\\
|
||
b = torch.zeros(3, 4) \emph{\# 3×4 全零矩阵}\\
|
||
c = torch.randn(3, 4) \emph{\# 3×4 随机矩阵}\\
|
||
\strut \\
|
||
\emph{\# 张量运算}\\
|
||
x = torch.tensor({[}1.0, 2.0, 3.0{]})\\
|
||
y = torch.tensor({[}4.0, 5.0, 6.0{]})\\
|
||
print(x + y) \emph{\# 加法}\\
|
||
print(torch.dot(x, y)) \emph{\# 点积}\\
|
||
\strut \\
|
||
\emph{\# NumPy互转}\\
|
||
\textbf{import} numpy \textbf{as} np\\
|
||
arr = np.array({[}1, 2, 3{]})\\
|
||
tensor = torch.from\_numpy(arr) \emph{\# NumPy → Tensor}\\
|
||
back = tensor.numpy() \emph{\# Tensor → NumPy}
|
||
|
||
\paragraph{自动求导(Autograd)}
|
||
|
||
PyTorch的autograd模块可以自动计算梯度,是训练神经网络的核心。
|
||
|
||
\emph{\# 创建需要梯度的张量}\\
|
||
x = torch.tensor({[}2.0{]}, requires\_grad=True)\\
|
||
\strut \\
|
||
\emph{\# 前向计算}\\
|
||
y = x ** 2 + 3 * x + 1 \emph{\# y = x² + 3x + 1}\\
|
||
\strut \\
|
||
\emph{\# 反向传播,自动计算 dy/dx}\\
|
||
y.backward()\\
|
||
print(f"dy/dx = \{x.grad\}") \emph{\# 应为 2x + 3 = 7(x=2时)}
|
||
|
||
\paragraph{完整示例:手写数字识别}
|
||
|
||
以下是一个完整的神经网络训练流程,使用经典的MNIST数据集。
|
||
|
||
\textbf{import} torch
|
||
|
||
\textbf{import} torch.nn \textbf{as} nn
|
||
|
||
\textbf{import} torch.optim \textbf{as} optim
|
||
|
||
\textbf{from} torchvision \textbf{import} datasets, transforms
|
||
|
||
|
||
|
||
\emph{\# 1. 数据准备}
|
||
|
||
transform = transforms.Compose({[}
|
||
|
||
\begin{lstlisting}
|
||
transforms.ToTensor(), # 转为张量
|
||
\end{lstlisting}
|
||
|
||
transforms.Normalize((0.1307,), (0.3081,)) \emph{\# 标准化}
|
||
|
||
{]})
|
||
|
||
train\_dataset = datasets.MNIST(\textquotesingle./data\textquotesingle,
|
||
train=True, download=True, transform=transform)
|
||
|
||
test\_dataset = datasets.MNIST(\textquotesingle./data\textquotesingle,
|
||
train=False, transform=transform)
|
||
|
||
train\_loader = torch.utils.data.DataLoader(train\_dataset,
|
||
batch\_size=64, shuffle=True)
|
||
|
||
test\_loader = torch.utils.data.DataLoader(test\_dataset,
|
||
batch\_size=1000)
|
||
|
||
\emph{\# 2. 定义模型}
|
||
|
||
\textbf{class} Net(nn.Module):
|
||
|
||
\textbf{def} \_\_init\_\_(self):
|
||
|
||
super().\_\_init\_\_()
|
||
|
||
self.fc1 = nn.Linear(28 * 28, 128) \emph{\# 输入层 → 隐藏层}
|
||
|
||
self.fc2 = nn.Linear(128, 64) \emph{\# 隐藏层 → 隐藏层}
|
||
|
||
self.fc3 = nn.Linear(64, 10) \emph{\# 隐藏层 → 输出层(10个数字)}
|
||
|
||
\textbf{def} forward(self, x):
|
||
|
||
\begin{lstlisting}
|
||
x = x.view(-1, 28 * 28) # 展平图像
|
||
\end{lstlisting}
|
||
|
||
x = torch.relu(self.fc1(x)) \emph{\# ReLU激活}
|
||
|
||
x = torch.relu(self.fc2(x))
|
||
|
||
\begin{lstlisting}
|
||
x = self.fc3(x) # 输出层不加激活
|
||
\end{lstlisting}
|
||
|
||
\textbf{return} x
|
||
|
||
model = Net()
|
||
|
||
\emph{\# 3. 定义损失函数和优化器}
|
||
|
||
criterion = nn.CrossEntropyLoss()
|
||
|
||
optimizer = optim.Adam(model.parameters(), lr=0.001)
|
||
|
||
\emph{\# 4. 训练}
|
||
|
||
\textbf{for} epoch \textbf{in} range(5):
|
||
|
||
model.train()
|
||
|
||
total\_loss = 0
|
||
|
||
\textbf{for} batch\_x, batch\_y \textbf{in} train\_loader:
|
||
|
||
\begin{lstlisting}
|
||
optimizer.zero_grad() # 清零梯度
|
||
output = model(batch_x) # 前向传播
|
||
\end{lstlisting}
|
||
|
||
loss = criterion(output, batch\_y) \emph{\# 计算损失}
|
||
|
||
\begin{lstlisting}
|
||
loss.backward() # 反向传播
|
||
optimizer.step() # 更新参数
|
||
\end{lstlisting}
|
||
|
||
total\_loss += loss.item()
|
||
|
||
print(f"Epoch \{epoch+1\}, Loss:
|
||
\{total\_loss/len(train\_loader):.4f\}")
|
||
|
||
\begin{lstlisting}
|
||
# 5. 测试
|
||
\end{lstlisting}
|
||
|
||
model.eval()
|
||
|
||
correct = 0
|
||
|
||
total = 0
|
||
|
||
\textbf{with} torch.no\_grad():
|
||
|
||
\textbf{for} batch\_x, batch\_y \textbf{in} test\_loader:
|
||
|
||
output = model(batch\_x)
|
||
|
||
\_, predicted = torch.max(output, 1)
|
||
|
||
total += batch\_y.size(0)
|
||
|
||
correct += (predicted == batch\_y).sum()
|
||
|
||
\begin{lstlisting}
|
||
print(f"\n测试准确率: {correct/total:.2%}")
|
||
\end{lstlisting}
|
||
|
||
\subsubsection{使用预训练模型}
|
||
|
||
在实际应用中,通常不需要从头训练模型,而是使用预训练模型进行微调或直接推理。
|
||
|
||
使用Hugging Face Transformers
|
||
|
||
\emph{\# 安装: pip install transformers}\\
|
||
\strut \\
|
||
\textbf{from} transformers \textbf{import} pipeline\\
|
||
\strut \\
|
||
\emph{\# 文本分类(情感分析)}\\
|
||
classifier = pipeline("sentiment-analysis")\\
|
||
result = classifier("This design is amazing!")\\
|
||
print(result)\\
|
||
\emph{\# {[}\{\textquotesingle label\textquotesingle:
|
||
\textquotesingle POSITIVE\textquotesingle,
|
||
\textquotesingle score\textquotesingle: 0.9998\}{]}}\\
|
||
\strut \\
|
||
\emph{\# 图像分类}\\
|
||
image\_classifier = pipeline("image-classification")\\
|
||
result = image\_classifier("building.jpg")\\
|
||
print(result)\\
|
||
\emph{\# {[}\{\textquotesingle score\textquotesingle: 0.92,
|
||
\textquotesingle label\textquotesingle:
|
||
\textquotesingle palace\textquotesingle\}, ...{]}}
|
||
|
||
\paragraph{使用Ultralytics YOLO}
|
||
|
||
\emph{\# 安装: pip install ultralytics}
|
||
|
||
\textbf{from} ultralytics \textbf{import} YOLO
|
||
|
||
\emph{\# 加载预训练模型}
|
||
|
||
model = YOLO("yolov8n.pt")
|
||
|
||
\emph{\# 目标检测}
|
||
|
||
results = model("street\_photo.jpg")
|
||
|
||
\emph{\# 查看结果}
|
||
|
||
\textbf{for} result \textbf{in} results:
|
||
|
||
boxes = result.boxes
|
||
|
||
\textbf{for} box \textbf{in} boxes:
|
||
|
||
cls = int(box.cls{[}0{]})
|
||
|
||
conf = float(box.conf{[}0{]})
|
||
|
||
label = model.names{[}cls{]}
|
||
|
||
\begin{lstlisting}
|
||
print(f"检测到: {label}, 置信度: {conf:.2f}")
|
||
\end{lstlisting}
|
||
|
||
\subsubsection{GPU加速}
|
||
|
||
深度学习训练在GPU上可以快数十倍。PyTorch的GPU使用非常简洁:
|
||
|
||
\emph{\# 检查GPU是否可用}\\
|
||
device = torch.device("cuda" \textbf{if} torch.cuda.is\_available()
|
||
\textbf{else} "cpu")\\
|
||
print(f"使用设备: \{device\}")\\
|
||
\strut \\
|
||
\emph{\# 将模型和数据移动到GPU}\\
|
||
model = Net().to(device)\\
|
||
\strut \\
|
||
\emph{\# 训练时,数据也需要移到GPU}\\
|
||
\textbf{for} batch\_x, batch\_y \textbf{in} train\_loader:\\
|
||
batch\_x = batch\_x.to(device)\\
|
||
batch\_y = batch\_y.to(device)\\
|
||
\emph{\# ... 后续训练代码不变}
|
||
|
||
在没有本地GPU的情况下,可以使用
|
||
\href{https://colab.research.google.com/}{Google Colab}
|
||
免费使用云端GPU运行上述代码。
|
||
|
||
\subsubsection{学习路径建议}
|
||
|
||
{\def\LTcaptype{none} % do not increment counter
|
||
\begin{longtable}[]{@{}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.1327}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.2478}}
|
||
>{\raggedright\arraybackslash}p{(\linewidth - 4\tabcolsep) * \real{0.6018}}@{}}
|
||
\toprule\noalign{}
|
||
\begin{minipage}[b]{\linewidth}\raggedright
|
||
阶段
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
内容
|
||
\end{minipage} & \begin{minipage}[b]{\linewidth}\raggedright
|
||
推荐资源
|
||
\end{minipage} \\
|
||
\midrule\noalign{}
|
||
\endhead
|
||
\bottomrule\noalign{}
|
||
\endlastfoot
|
||
入门 & Python基础 + NumPy/Pandas & 本附录 §3 \\
|
||
机器学习 & scikit-learn实践 &
|
||
\href{https://scikit-learn.org/stable/tutorial/}{scikit-learn官方教程} \\
|
||
深度学习基础 & PyTorch入门 + MLP &
|
||
\href{https://pytorch.org/tutorials/}{PyTorch官方教程} \\
|
||
计算机视觉 & CNN + 图像分类/检测 & 本书第三篇 + CS231n \\
|
||
自然语言处理 & Transformer + LLM & 本书第四篇 + CS224n \\
|
||
生成式AI & Diffusion + AIGC工具 & 本书第五篇 \\
|
||
前沿探索 & Agent + 具身智能 & 本书第六篇 \\
|
||
\end{longtable}
|
||
}
|