在 Bun 环境下用 TypeScript 编写一个网站爬虫是一个很有趣的项目。Bun 是一个新兴的 JavaScript 运行时,兼具快速和现代化的特点。下面是一个简单的网站爬虫实例,展示了如何在 Bun 中使用 TypeScript 进行网页抓取。
首先,确保你已经安装了 Bun。如果没有,请按照官方文档进行安装。
步骤一:设置项目
-
创建一个新的项目目录并初始化 Bun 项目:
mkdir my-web-crawler cd my-web-crawler bun init -
安装 TypeScript 和相关类型定义:
bun add -d typescript @types/node -
初始化 TypeScript 配置文件:
npx tsc --init
步骤二:安装所需依赖
为了进行 HTTP 请求和解析 HTML,我们需要安装一些库:
bun add axios cheerio步骤三:编写爬虫代码
创建一个 src 目录并在其中创建一个 crawler.ts 文件:
// src/crawler.ts
import axios from 'axios';
import cheerio from 'cheerio';
const url = 'https://example.com'; // 你要抓取的网站URL
async function fetchHTML(url: string): Promise<string> {
const { data } = await axios.get(url);
return data;
}
async function parseHTML(html: string): Promise<void> {
const $ = cheerio.load(html);
$('a').each((index, element) => {
const link = $(element).attr('href');
if (link) {
console.log(`Found link: ${link}`);
}
});
}
async function main() {
try {
const html = await fetchHTML(url);
await parseHTML(html);
} catch (error) {
console.error('Error occurred:', error);
}
}
main();步骤四:配置 TypeScript
编辑 tsconfig.json,确保编译目标和模块解析适用于 Bun 环境:
{
"compilerOptions": {
"target": "es2020",
"module": "esnext",
"moduleResolution": "node",
"outDir": "./dist",
"rootDir": "./src",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true
},
"include": ["src"]
}步骤五:编译并运行爬虫
编译 TypeScript 代码:
bun tsc运行编译后的 JavaScript 代码:
bun run dist/crawler.js代码解释
fetchHTML函数使用axios库发送 HTTP GET 请求并返回 HTML 字符串。parseHTML函数使用cheerio库解析 HTML 并提取所有链接。main函数协调上述两个步骤,并处理可能的错误。
通过这个示例,你可以了解如何在 Bun 环境下使用 TypeScript 编写一个简单的网页爬虫。根据需要,你可以扩展这个爬虫,增加更多的功能,比如处理不同的页面、遵循机器人协议(robots. Txt)、并发抓取等。