init: KG_ICH 项目初始化
- data/: 非遗地理编码数据(GIS shapefile + CSV) - dofile/kg_project/: 知识图谱构建代码(纳入主仓库) - dofile/visulization/: 可视化数据与路线图 - officefile/: 文献、草稿、bib 文档 - officefile/latex/: Overleaf 同步目录(独立管理,不纳入) - output/: 输出目录 - logs/: 日志目录
This commit is contained in:
+29
@@ -0,0 +1,29 @@
|
|||||||
|
# === 工具与缓存 ===
|
||||||
|
.claude/
|
||||||
|
.playwright-mcp/
|
||||||
|
*.log
|
||||||
|
|
||||||
|
# === Python 虚拟环境 ===
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
|
||||||
|
# === Obsidian & Pandoc ===
|
||||||
|
.obsidian/
|
||||||
|
.pandoc/
|
||||||
|
|
||||||
|
# === Overleaf Git(仅 latex/ 同步)===
|
||||||
|
officefile/latex/.git/
|
||||||
|
|
||||||
|
# === IDE / OS ===
|
||||||
|
.vscode/
|
||||||
|
.idea/
|
||||||
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
|
||||||
|
# === 旧的嵌套 Git 历史 ===
|
||||||
|
.git-old-local/
|
||||||
|
|
||||||
|
# === 大型下载文件 ===
|
||||||
|
downloads/
|
||||||
Binary file not shown.
Binary file not shown.
@@ -0,0 +1 @@
|
|||||||
|
GEOGCS["GCS_WGS_1984",DATUM["D_WGS_1984",SPHEROID["WGS_1984",6378137.0,298.257223563]],PRIMEM["Greenwich",0.0],UNIT["Degree",0.0174532925199433]]
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,269 @@
|
|||||||
|
总序号,项目保护单位,bd09_lat,bd09_lon,geocode_status,wgs84_lon,wgs84_lat
|
||||||
|
1,黑龙江省艺术研究所,45.75582106,126.6438083,success,126.6312812,45.74801106
|
||||||
|
2,齐齐哈尔市梅里斯达斡尔族区文化馆,47.31554957,123.7595409,success,123.7461629,47.3074024
|
||||||
|
3,黑河市艺术研究所,50.25127231,127.5354899,success,127.5217199,50.24367861
|
||||||
|
4,黑龙江省群众艺术馆,45.77151098,126.6480975,success,126.6355371,45.76377588
|
||||||
|
5,伊春市戏曲剧院,47.73331846,128.8475464,success,128.8339473,47.72523303
|
||||||
|
6,方正县文化馆,45.85775844,128.8356337,success,128.8222747,45.85006293
|
||||||
|
7,东宁县文化馆,44.0673227,131.1294821,success,131.1152579,44.05891173
|
||||||
|
8,七台河市非物质文化遗产保护中心,45.78213277,131.0196313,success,131.005397,45.77381386
|
||||||
|
9,肇源县非物质文化遗产保护中心,45.52415291,125.0845726,success,125.0711585,45.51602736
|
||||||
|
10,双城市非物质文化遗产保护中心,45.38811152,126.3196231,success,126.3067329,45.38019948
|
||||||
|
11,五常市文化馆,44.94382191,127.1776152,success,127.1647082,44.93562481
|
||||||
|
12,海伦市人民艺术剧院,47.45690384,126.9365086,success,126.9236728,47.44873835
|
||||||
|
13,绥化市北林区文工团,46.63862198,126.97987,success,126.9671481,46.6304274
|
||||||
|
14,肇州县非物质文化遗产保护中心,45.71009294,125.2906327,success,125.2775591,45.70163614
|
||||||
|
15,黑龙江省龙江剧艺术中心,47.18896666,124.8840662,success,124.8709297,47.18117849
|
||||||
|
16,双城市文化馆,45.38916578,126.320916,success,126.3080285,45.38123422
|
||||||
|
17,绥化市绥棱县文化馆,47.24251579,127.1205151,success,127.1070965,47.23415207
|
||||||
|
18,北安市评剧团,48.24741953,126.4973797,success,126.4843479,48.23891986
|
||||||
|
19,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
20,鹤岗市群众艺术馆,47.34926432,130.2840214,success,130.2696228,47.34108801
|
||||||
|
21,黑龙江省曲艺团有限公司,45.76809,126.66036,success,126.6477152,45.76047629
|
||||||
|
22,桦川县文化馆,47.03658466,130.7293638,success,130.7149988,47.02857188
|
||||||
|
23,黑龙江省杂技团,45.73284876,126.6106653,success,126.5983848,45.72455485
|
||||||
|
24,齐齐哈尔市马戏团,47.36444108,123.9339047,success,123.9208401,47.35644797
|
||||||
|
25,大庆市杜尔伯特蒙古族自治县博物馆,46.86876776,124.4493588,success,124.4361789,46.86024699
|
||||||
|
26,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
27,五常市文化馆,44.94382191,127.1776152,success,127.1647082,44.93562481
|
||||||
|
28,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
29,宁安市文化体育局,44.34538409,129.476914,success,129.4632135,44.33677404
|
||||||
|
30,双鸭山市饶河县文化馆,46.81581498,134.0176875,success,134.0037288,46.80749762
|
||||||
|
31,黑龙江省传统武学研究会,45.74792984,126.6696528,success,126.6569785,45.74030715
|
||||||
|
32,黑龙江省归国华侨联合会,45.74338857,126.6775305,success,126.664857,45.73569396
|
||||||
|
33,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
34,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
35,杜尔伯特蒙古族自治县文化体育局,46.87300908,124.4529011,success,124.4397277,46.86451077
|
||||||
|
36,黑龙江省艺术研究所,45.75582106,126.6438083,success,126.6312812,45.74801106
|
||||||
|
37,五常市非物质文化遗产保护中心,44.93784286,127.1735288,success,127.1605796,44.92971861
|
||||||
|
38,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
39,阿城区民间文艺家协会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
40,宁安市文化体育局,44.34538409,129.476914,success,129.4632135,44.33677404
|
||||||
|
41,富锦市群众艺术馆,47.26608314,132.0412614,success,132.0268239,47.25766497
|
||||||
|
42,东宁县文化馆,44.0673227,131.1294821,success,131.1152579,44.05891173
|
||||||
|
43,双城市文化馆,45.38916578,126.320916,success,126.3080285,45.38123422
|
||||||
|
44,依兰县文化馆,46.33126029,129.5745197,success,129.560623,46.32343594
|
||||||
|
45,富裕县文化馆,47.80583728,124.4834558,success,124.4700126,47.79765568
|
||||||
|
46,兰西县文化馆,46.26732644,126.2978939,success,126.2848964,46.25969498
|
||||||
|
47,富裕县文化馆,47.80583728,124.4834558,success,124.4700126,47.79765568
|
||||||
|
48,牡丹江市非物质文化遗产保护协会,44.58185651,129.6111013,success,129.597687,44.57353907
|
||||||
|
49,肇州县非物质文化遗产保护中心,45.71009294,125.2906327,success,125.2775591,45.70163614
|
||||||
|
50,密山市文化馆,45.55448623,131.8864766,success,131.8728149,45.5459851
|
||||||
|
51,五大连池药泉民俗研究会,48.66257739,126.1514718,success,126.1379422,48.6542552
|
||||||
|
52,讷河市非物质文化遗产保护中心,48.47252804,124.8891678,success,124.8758643,48.46479829
|
||||||
|
53,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
54,齐齐哈尔市铁锋区文化体育中心,47.34701886,123.9844157,success,123.9713027,47.3386483
|
||||||
|
55,肇源县文化活动中心,45.52415291,125.0845726,success,125.0711585,45.51602736
|
||||||
|
56,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
57,哈尔滨市非物质文化遗产保护中心,45.771343,126.648475,success,126.6359114,45.76361335
|
||||||
|
58,佳木斯市郊区非物质文化遗产保护中心,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
59,佳木斯市群众艺术馆,46.825782,130.369945,success,130.3554282,46.81771885
|
||||||
|
60,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
61,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
62,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
63,大庆市肇州县文化馆,45.71622277,125.2790967,success,125.2660297,45.70779699
|
||||||
|
64,伊春市艺术研究室,47.73331846,128.8475464,success,128.8339473,47.72523303
|
||||||
|
65,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
66,大庆市杜尔伯特蒙古族自治县博物馆,46.86876776,124.4493588,success,124.4361789,46.86024699
|
||||||
|
67,双鸭山市饶河县文化馆,46.81581498,134.0176875,success,134.0037288,46.80749762
|
||||||
|
68,佳木斯市郊区非物质文化遗产保护中心,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
69,杜尔伯特蒙古族自治县文化体育局,46.87300908,124.4529011,success,124.4397277,46.86451077
|
||||||
|
70,哈尔滨市阿城区音乐家协会,45.54991475,126.9827822,success,126.9702047,45.54158713
|
||||||
|
71,齐齐哈尔市梅里斯达斡尔族区文化馆,47.31554957,123.7595409,success,123.7461629,47.3074024
|
||||||
|
72,黑河市爱辉区文化馆,50.25261678,127.5213824,success,127.5074536,50.2452366
|
||||||
|
73,佳木斯市群众艺术馆,47.26608314,132.0412614,success,132.0268239,47.25766497
|
||||||
|
74,牡丹江市朝鲜民族艺术馆,44.58301369,129.6111989,success,129.5977853,44.57469723
|
||||||
|
75,牡丹江市朝鲜民族艺术馆,44.58301369,129.6111989,success,129.5977853,44.57469723
|
||||||
|
76,宁安市非物质文化遗产保护协会,44.34698358,129.489368,success,129.4757255,44.3383656
|
||||||
|
77,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
78,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
79,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
80,林甸县文化馆,47.18968604,124.8832976,success,124.8701555,47.1819085
|
||||||
|
81,七台河市茄子河区文化馆,45.79123818,131.0744806,success,131.06004,45.78282072
|
||||||
|
82,齐齐哈尔市梅里斯达斡尔族区文化馆,47.31554957,123.7595409,success,123.7461629,47.3074024
|
||||||
|
83,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
84,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
85,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
86,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
87,双鸭山市饶河县文化馆,46.81581498,134.0176875,success,134.0037288,46.80749762
|
||||||
|
88,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
89,泰来县文化馆,46.39782929,123.4236263,success,123.4101272,46.39017044
|
||||||
|
90,哈尔滨市阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
91,讷河市文化馆,48.47252804,124.8891678,success,124.8758643,48.46479829
|
||||||
|
92,同江市群众艺术馆,47.64798068,132.5175095,success,132.5034909,47.63953006
|
||||||
|
93,佳木斯市郊区非物质文化遗产保护中心,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
94,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
95,牡丹江市朝鲜民族艺术馆,44.58301369,129.6111989,success,129.5977853,44.57469723
|
||||||
|
96,嘉荫县群众艺术馆,48.89498347,130.4105555,success,130.3955959,48.88668004
|
||||||
|
97,嘉荫县文化馆,48.89407887,130.4071028,success,130.3921184,48.88584037
|
||||||
|
98,牡丹江市朝鲜民族艺术馆,44.58301369,129.6111989,success,129.5977853,44.57469723
|
||||||
|
99,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
100,牡丹江市朝鲜民族艺术馆,44.58301369,129.6111989,success,129.5977853,44.57469723
|
||||||
|
101,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
102,哈尔滨市朝鲜民族艺术馆,45.76766336,126.6190306,success,126.6067046,45.75944138
|
||||||
|
103,东宁县文化馆,44.0673227,131.1294821,success,131.1152579,44.05891173
|
||||||
|
104,甘南县文化馆,47.87611042,123.7061596,success,123.6928614,47.86778014
|
||||||
|
105,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
106,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
107,五常市文化馆,44.94382191,127.1776152,success,127.1647082,44.93562481
|
||||||
|
108,宁安市文化体育局,44.34538409,129.476914,success,129.4632135,44.33677404
|
||||||
|
109,大庆市群众艺术馆,46.59363318,125.1086576,success,125.0949471,46.58589537
|
||||||
|
110,青冈县文化馆,46.68862512,126.1085948,success,126.0953877,46.68024011
|
||||||
|
111,鸡西市群众艺术馆,45.30677605,130.9438395,success,130.9300798,45.29843869
|
||||||
|
112,庆安县文化馆,46.88574447,127.5146122,success,127.5012918,46.87802058
|
||||||
|
113,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
114,克山县文化馆,48.04038692,125.8798207,success,125.8667532,48.03198628
|
||||||
|
115,大庆市让胡路区百湖城书艺馆,46.64173096,124.8669754,success,124.8538876,46.63433587
|
||||||
|
116,伊春市艺术研究室,47.73331846,128.8475464,success,128.8339473,47.72523303
|
||||||
|
117,望奎县文化馆,46.84411964,126.4760837,success,126.463194,46.83564973
|
||||||
|
118,哈尔滨儿童艺术剧院,45.79403096,126.6624385,success,126.6498027,45.78643997
|
||||||
|
119,双城市皮影艺术团,45.38811152,126.3196231,success,126.3067329,45.38019948
|
||||||
|
120,黑龙江省评剧艺术中心,45.78924165,126.6522311,success,126.6396516,45.78157271
|
||||||
|
121,佳木斯市评剧团,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
122,黑龙江省京剧院,45.75884697,126.6553175,success,126.6426998,45.75119447
|
||||||
|
123,哈尔滨市方正县文化馆,45.85775844,128.8356337,success,128.8222747,45.85006293
|
||||||
|
124,海伦市文化馆,47.45690384,126.9365086,success,126.9236728,47.44873835
|
||||||
|
125,绥化市兰西县文化馆,46.26732644,126.2978939,success,126.2848964,46.25969498
|
||||||
|
126,呼玛县文化馆,51.73086307,126.6704831,success,126.6566811,51.72372518
|
||||||
|
127,佳木斯市民间艺术家协会,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
128,北方民俗文化艺术研究中心,46.599149,125.175051,success,125.1617004,46.59078096
|
||||||
|
129,齐齐哈尔市龙沙区文化馆,47.32357698,123.9643762,success,123.9513764,47.31510979
|
||||||
|
130,安达市秀英民间艺术剪纸画廊,46.45796584,125.3141247,success,125.3007241,46.44995417
|
||||||
|
131,北安市评剧团,48.24741953,126.4973797,success,126.4843479,48.23891986
|
||||||
|
132,五大连池市文物管理所,48.5145725,126.200389,success,126.1868924,48.50670807
|
||||||
|
133,黑龙江禹舜文化艺术研究院,45.76049549,126.6534938,success,126.6408885,45.75282515
|
||||||
|
134,宁安市文化体育局,44.34538409,129.476914,success,129.4632135,44.33677404
|
||||||
|
135,黑河市爱辉区文化馆,50.25261678,127.5213824,success,127.5074536,50.2452366
|
||||||
|
136,佳木斯市群众艺术馆,47.26608314,132.0412614,success,132.0268239,47.25766497
|
||||||
|
137,哈尔滨市群众艺术馆,45.77151098,126.6480975,success,126.6355371,45.76377588
|
||||||
|
138,绥棱县文化馆,47.24251579,127.1205151,success,127.1070965,47.23415207
|
||||||
|
139,北安市评剧团,48.24741953,126.4973797,success,126.4843479,48.23891986
|
||||||
|
140,大兴安岭塔河县群众艺术馆,52.34030508,124.7165125,success,124.7022813,52.33289247
|
||||||
|
141,克东县北方满绣艺术研究所,48.04824416,126.2553867,success,126.2423338,48.03968479
|
||||||
|
142,齐齐哈尔市群众艺术馆,47.34737244,123.9514263,success,123.938436,47.33903986
|
||||||
|
143,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
144,北安市评剧团,48.24741953,126.4973797,success,126.4843479,48.23891986
|
||||||
|
145,宁安市非物质文化遗产保护协会,44.34698358,129.489368,success,129.4757255,44.3383656
|
||||||
|
146,大庆市肇源县文化活动中心,45.52415291,125.0845726,success,125.0711585,45.51602736
|
||||||
|
147,哈尔滨市阿城区民间文艺家协会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
148,安达市文化体育局,46.45796584,125.3141247,success,125.3007241,46.44995417
|
||||||
|
149,史作玺糖艺面塑创作室,45.713621,126.669629,success,126.6569888,45.70600004
|
||||||
|
150,哈尔滨市阿城区民间文艺家协会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
151,哈尔滨市特种工艺美术厂,45.76574557,126.6797049,success,126.667029,45.75802654
|
||||||
|
152,哈尔滨市工艺美术有限责任公司,45.76574557,126.6797049,success,126.667029,45.75802654
|
||||||
|
153,同江市群众艺术馆,47.64798068,132.5175095,success,132.5034909,47.63953006
|
||||||
|
154,齐齐哈尔市群众艺术馆,47.34737244,123.9514263,success,123.938436,47.33903986
|
||||||
|
155,齐齐哈尔市群众艺术馆,47.34737244,123.9514263,success,123.938436,47.33903986
|
||||||
|
156,伊春市艺术研究室,47.73331846,128.8475464,success,128.8339473,47.72523303
|
||||||
|
157,汤原县文化馆,46.7370677,129.9221257,success,129.9080552,46.729408
|
||||||
|
158,大庆市非物质文化遗产保护中心,46.58901729,125.1623722,success,125.1490055,46.58060067
|
||||||
|
159,绥棱县文化馆,47.24251579,127.1205151,success,127.1070965,47.23415207
|
||||||
|
160,黑龙江国粹戏剧艺术博物馆,45.77145132,126.6473829,success,126.6348284,45.76370546
|
||||||
|
161,伊春市乌马河区文化馆,48.89407887,130.4071028,success,130.3921184,48.88584037
|
||||||
|
162,黑龙江龙广之声文化传播有限责任公司,45.753032,126.687176,success,126.6745198,45.74517987
|
||||||
|
163,黑龙江省民族博物馆,45.78172497,126.6814856,success,126.6688227,45.77398752
|
||||||
|
164,黑河市爱辉区文物管理所,50.25550311,127.515815,success,127.5018354,50.24817418
|
||||||
|
165,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
166,双鸭山市饶河县文化馆,46.81581498,134.0176875,success,134.0037288,46.80749762
|
||||||
|
167,黑河市非物质文化遗产保护中心,50.2480745,127.5204212,success,127.5064827,50.24070154
|
||||||
|
168,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
169,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
170,哈尔滨日炎公司,45.80882583,126.5416151,success,126.5289594,45.80121771
|
||||||
|
171,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
172,双城市文化馆,45.38916578,126.320916,success,126.3080285,45.38123422
|
||||||
|
173,哈尔滨秋林食品厂,45.67775619,126.616705,success,126.6044307,45.66949712
|
||||||
|
174,哈尔滨秋林里道斯食品有限责任公司,45.63337857,126.817157,success,126.8045267,45.62528875
|
||||||
|
175,哈尔滨大众肉联集团有限公司,45.53244496,126.5343041,success,126.5216586,45.52468904
|
||||||
|
176,哈尔滨老鼎丰食品有限公司,45.91133879,126.5535234,success,126.5408838,45.9037732
|
||||||
|
177,哈尔滨老都一处餐饮有限责任公司,45.80882583,126.5416151,success,126.5289594,45.80121771
|
||||||
|
178,哈尔滨市老厨家道台食府,45.70992736,126.6667568,success,126.6541261,45.70231692
|
||||||
|
179,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
180,克东县腐乳产业协会,48.0400769,126.2541175,success,126.2410659,48.03150272
|
||||||
|
181,黑龙江省调味品工业协会,45.74792984,126.6696528,success,126.6569785,45.74030715
|
||||||
|
182,绥棱县文化馆,47.24251579,127.1205151,success,127.1070965,47.23415207
|
||||||
|
183,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
184,伊春市艺术研究室,47.73331846,128.8475464,success,128.8339473,47.72523303
|
||||||
|
185,双城市白酒工艺研究会,45.38811152,126.3196231,success,126.3067329,45.38019948
|
||||||
|
186,黑龙江北大仓集团有限公司,45.751576,126.676666,success,126.6639869,45.74389303
|
||||||
|
187,黑龙江省富裕老窖酒业有限责任公司,45.7401911,126.6824096,success,126.6697471,45.73242424
|
||||||
|
188,庆安县文化馆,46.88574447,127.5146122,success,127.5012918,46.87802058
|
||||||
|
189,黑龙江省玉泉酒业有限责任公司,45.42151265,127.167268,success,127.1541858,45.41341576
|
||||||
|
190,双城市关东酒业有限公司,45.38811152,126.3196231,success,126.3067329,45.38019948
|
||||||
|
191,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
192,黑龙江省望奎高贤酒业有限公司,47.06606258,126.6818311,success,126.6689398,47.0582359
|
||||||
|
193,穆棱市文化馆,44.92398753,130.5362661,success,130.5225358,44.91590669
|
||||||
|
194,七台河市群众艺术馆,45.78174652,131.0466262,success,131.0323887,45.77302433
|
||||||
|
195,牡丹江市群众艺术馆、渤海镇文化站,44.12687529,129.1759754,success,129.1623219,44.11882468
|
||||||
|
196,穆棱市文化馆,44.92398753,130.5362661,success,130.5225358,44.91590669
|
||||||
|
197,饶河县非物质文化遗产保护中心,46.80418274,134.0204689,success,134.006533,46.79580832
|
||||||
|
198,绥棱县文化馆,47.24251579,127.1205151,success,127.1070965,47.23415207
|
||||||
|
199,黑龙江省兴勃工艺品有限公司,45.74792984,126.6696528,success,126.6569785,45.74030715
|
||||||
|
200,李氏黑陶文化艺术研究所,45.631599,126.52717,success,126.5145109,45.62373057
|
||||||
|
201,依安县文化馆,47.91489855,125.3204008,success,125.3066585,47.90666647
|
||||||
|
202,绥化市泥河陶文化艺术有限公司,47.3142979,128.007295,success,127.9939082,47.30614389
|
||||||
|
203,五大连池市文物管理所,48.5145725,126.200389,success,126.1868924,48.50670807
|
||||||
|
204,五常市文化馆,44.94382191,127.1776152,success,127.1647082,44.93562481
|
||||||
|
205,哈尔滨市庆成火匏产品有限公司,45.80882583,126.5416151,success,126.5289594,45.80121771
|
||||||
|
206,七台河市英雄工艺刀剑锻造有限公司,45.77630032,131.0115446,success,130.997287,45.76816744
|
||||||
|
207,阿城区民间工艺美术家协会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
208,黑龙江省文化艺术发展中心,45.30842136,130.9898146,success,130.9756641,45.30060499
|
||||||
|
209,哈尔滨市道里区文化馆,45.77853084,126.6146262,success,126.6023267,45.77027229
|
||||||
|
210,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
211,黑河市爱辉区文化馆,50.25261678,127.5213824,success,127.5074536,50.2452366
|
||||||
|
212,呼玛县文化馆,51.73086307,126.6704831,success,126.6566811,51.72372518
|
||||||
|
213,宾县文化馆,45.76361288,127.4922455,success,127.479025,45.75572369
|
||||||
|
214,七台河华韵民族乐器制作有限公司,45.77630032,131.0115446,success,130.997287,45.76816744
|
||||||
|
215,三合缘根艺家具厂,47.707191,128.958087,success,128.9443187,47.69924788
|
||||||
|
216,七台河市经济开发区郝家木艺坊,45.77630032,131.0115446,success,130.997287,45.76816744
|
||||||
|
217,抚远县文化馆,48.37363711,134.3129708,success,134.2987716,48.36515205
|
||||||
|
218,哈尔滨洪一轩文化传媒有限公司,45.8815795,126.5441568,success,126.5314593,45.87399685
|
||||||
|
219,黑龙江省装裱艺术研究会,45.74792984,126.6696528,success,126.6569785,45.74030715
|
||||||
|
220,黑龙江省图书馆,45.36780947,126.3090899,success,126.2962097,45.36000527
|
||||||
|
221,黑河市爱辉区文化馆,50.25261678,127.5213824,success,127.5074536,50.2452366
|
||||||
|
222,黑河市五大连池风景区非物质文化遗产保护中心,48.70259497,126.2714585,success,126.2582276,48.69441304
|
||||||
|
223,讷河市文化馆,48.47252804,124.8891678,success,124.8758643,48.46479829
|
||||||
|
224,牡丹江市朝鲜民族艺术馆,44.583022,129.611217,success,129.5978034,44.57470563
|
||||||
|
225,鸡西市朝鲜族艺术馆,45.30254683,130.9731084,success,130.9590567,45.29468866
|
||||||
|
226,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
227,杜尔伯特蒙古族自治县文化体育局,46.87300908,124.4529011,success,124.4397277,46.86451077
|
||||||
|
228,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
229,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
230,同江市群众艺术馆,47.64798068,132.5175095,success,132.5034909,47.63953006
|
||||||
|
231,兰西锡伯民族文化中心,46.25812558,126.2964168,success,126.2834192,46.25047767
|
||||||
|
232,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
233,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
234,阿城区民间文艺家协会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
235,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
236,齐齐哈尔市富拉尔基区文化馆,47.21590169,123.6458875,success,123.6327132,47.2080677
|
||||||
|
237,齐齐哈尔市梅里斯达斡尔族区文化馆,47.31554957,123.7595409,success,123.7461629,47.3074024
|
||||||
|
238,黑河市爱辉区文化馆,50.25261678,127.5213824,success,127.5074536,50.2452366
|
||||||
|
239,齐齐哈尔市富裕县文物管理所,47.77758936,124.4766437,success,124.4632343,47.76925486
|
||||||
|
240,大庆市肇源县文化活动中心,45.52415291,125.0845726,success,125.0711585,45.51602736
|
||||||
|
241,海林市非物质文化遗产保护协会,44.59987197,129.3874268,success,129.3735685,44.59140926
|
||||||
|
242,饶河县文化局,46.80853529,134.031975,success,134.0180638,46.79991581
|
||||||
|
243,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
244,宁安市非物质文化遗产保护协会,44.34698358,129.489368,success,129.4757255,44.3383656
|
||||||
|
245,齐齐哈尔市梅里斯达斡尔族区文化馆,47.31554957,123.7595409,success,123.7461629,47.3074024
|
||||||
|
246,木兰县文化馆,45.9484543,128.0610095,success,128.047872,45.93995843
|
||||||
|
247,兰西县文化馆,46.26732644,126.2978939,success,126.2848964,46.25969498
|
||||||
|
248,黑河市民族宗教事务局民族政策研究中心,50.25127231,127.5354899,success,127.5217199,50.24367861
|
||||||
|
249,大兴安岭地区群众艺术馆,50.42823532,124.138259,success,124.1240089,50.42089502
|
||||||
|
250,杜尔伯特蒙古族自治县文化体育局,46.87300908,124.4529011,success,124.4397277,46.86451077
|
||||||
|
251,大庆市杜尔伯特蒙古族自治县博物馆,46.86876776,124.4493588,success,124.4361789,46.86024699
|
||||||
|
252,同江市群众艺术馆,47.64798068,132.5175095,success,132.5034909,47.63953006
|
||||||
|
253,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
254,阿城区民俗学会,45.5542753,126.9643565,success,126.9519054,45.54586951
|
||||||
|
255,宾县文化馆,45.76361288,127.4922455,success,127.479025,45.75572369
|
||||||
|
256,黑龙江省民族博物馆,45.78172497,126.6814856,success,126.6688227,45.77398752
|
||||||
|
257,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
258,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
259,宁安市文化体育局,44.34538409,129.476914,success,129.4632135,44.33677404
|
||||||
|
260,阿城区满族联谊会,45.52847968,126.9746364,success,126.9621138,45.52006942
|
||||||
|
261,呼玛县文化馆,51.73086307,126.6704831,success,126.6566811,51.72372518
|
||||||
|
262,呼玛县文化馆,51.73086307,126.6704831,success,126.6566811,51.72372518
|
||||||
|
263,呼玛县文化馆,51.73086307,126.6704831,success,126.6566811,51.72372518
|
||||||
|
264,逊克县博物馆,49.56949144,128.4855846,success,128.4716087,49.56148024
|
||||||
|
265,宁安市文化馆,44.35308924,129.4761375,success,129.4624379,44.34449078
|
||||||
|
266,佳木斯市郊区非物质文化遗产保护中心,46.80568999,130.3273591,success,130.3132017,46.79707441
|
||||||
|
267,海林市文化馆,44.5780916,129.3953204,success,129.3813797,44.56976716
|
||||||
|
268,牡丹江市群众艺术馆,44.58853502,129.625767,success,129.6122832,44.58038211
|
||||||
|
@@ -0,0 +1,102 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
import pandas as pd
|
||||||
|
import json
|
||||||
|
import time
|
||||||
|
import os
|
||||||
|
from urllib.request import urlopen, quote
|
||||||
|
|
||||||
|
# 添加skill scripts到路径
|
||||||
|
skill_dir = r'C:\Users\xiaopeng\.claude\skills\geocoding-cn\scripts'
|
||||||
|
if skill_dir not in os.sys.path:
|
||||||
|
os.sys.path.insert(0, skill_dir)
|
||||||
|
|
||||||
|
from coordinate_transform import CoordinateTransformer
|
||||||
|
|
||||||
|
# 读取数据
|
||||||
|
input_file = r'E:\Project\2026_KG_ICH\data\黑龙江国家级和省级非遗名单-地理编码-最终.xlsx'
|
||||||
|
df = pd.read_excel(input_file)
|
||||||
|
|
||||||
|
# 百度地图API配置
|
||||||
|
AK = "L5SlQ1Kwmg6zaESvmc6RKG37yK2va7Ry"
|
||||||
|
BASE_URL = 'https://api.map.baidu.com/geocoding/v3/'
|
||||||
|
|
||||||
|
def geocode_with_retry(address):
|
||||||
|
"""尝试多次地理编码,使用不同的地址格式"""
|
||||||
|
attempts = [
|
||||||
|
address, # 原始地址
|
||||||
|
f"黑龙江省{address}", # 添加省名
|
||||||
|
f"{address}黑龙江", # 省名在后
|
||||||
|
f"中国黑龙江省{address}", # 添加国家
|
||||||
|
]
|
||||||
|
|
||||||
|
for attempt in attempts:
|
||||||
|
try:
|
||||||
|
encoded_address = quote(attempt)
|
||||||
|
url = f'{BASE_URL}?address={encoded_address}&output=json&ak={AK}'
|
||||||
|
req = urlopen(url, timeout=10)
|
||||||
|
response = req.read().decode()
|
||||||
|
result = json.loads(response)
|
||||||
|
|
||||||
|
if result['status'] == 0:
|
||||||
|
location = result['result']['location']
|
||||||
|
return location['lat'], location['lng'], attempt
|
||||||
|
except Exception as e:
|
||||||
|
continue
|
||||||
|
|
||||||
|
time.sleep(0.2)
|
||||||
|
|
||||||
|
return None, None, None
|
||||||
|
|
||||||
|
# 问题记录索引(0-based)
|
||||||
|
problematic_indices = [20, 127, 148, 161, 185, 199, 214, 58, 223]
|
||||||
|
|
||||||
|
print("=== 重新地理编码问题记录 ===\n")
|
||||||
|
|
||||||
|
success_count = 0
|
||||||
|
for idx in problematic_indices:
|
||||||
|
if idx >= len(df):
|
||||||
|
continue
|
||||||
|
|
||||||
|
row = df.iloc[idx]
|
||||||
|
original_address = str(row.iloc[1]) # 项目保护单位列
|
||||||
|
|
||||||
|
print(f"Row {idx+1}: {original_address}")
|
||||||
|
print(f" Old coords: {row['wgs84_lat']:.4f}N, {row['wgs84_lon']:.4f}E")
|
||||||
|
|
||||||
|
# 尝试重新地理编码
|
||||||
|
lat, lon, used_address = geocode_with_retry(original_address)
|
||||||
|
|
||||||
|
if lat and lon:
|
||||||
|
# 转换为WGS84
|
||||||
|
wgs_lon, wgs_lat = CoordinateTransformer.bd09_to_wgs84(lon, lat)
|
||||||
|
|
||||||
|
print(f" New BD09: {lat:.4f}N, {lon:.4f}E")
|
||||||
|
print(f" New WGS84: {wgs_lat:.4f}N, {wgs_lon:.4f}E")
|
||||||
|
print(f" Used address: {used_address}")
|
||||||
|
|
||||||
|
# 检查是否在合理范围内
|
||||||
|
if 43 <= wgs_lat <= 53 and 121 <= wgs_lon <= 135:
|
||||||
|
print(f" OK: Within Heilongjiang range")
|
||||||
|
# 更新数据
|
||||||
|
df.at[idx, 'bd09_lat'] = lat
|
||||||
|
df.at[idx, 'bd09_lon'] = lon
|
||||||
|
df.at[idx, 'wgs84_lat'] = wgs_lat
|
||||||
|
df.at[idx, 'wgs84_lon'] = wgs_lon
|
||||||
|
df.at[idx, 'geocode_status'] = 'success'
|
||||||
|
success_count += 1
|
||||||
|
else:
|
||||||
|
print(f" WARNING: Still outside range (43-53N, 121-135E)")
|
||||||
|
else:
|
||||||
|
print(f" FAILED: Geocoding failed")
|
||||||
|
|
||||||
|
print()
|
||||||
|
time.sleep(0.3)
|
||||||
|
|
||||||
|
# 保存更新后的文件
|
||||||
|
output_file = r'E:\Project\2026_KG_ICH\data\黑龙江国家级和省级非遗名单-地理编码-最终-修正.xlsx'
|
||||||
|
df.to_excel(output_file, index=False)
|
||||||
|
print(f"\n{'='*60}")
|
||||||
|
print(f"Correction complete!")
|
||||||
|
print(f"Successfully corrected: {success_count}/{len(problematic_indices)} records")
|
||||||
|
print(f"Saved to: {output_file}")
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,57 @@
|
|||||||
|
# Python
|
||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
*$py.class
|
||||||
|
*.so
|
||||||
|
.Python
|
||||||
|
|
||||||
|
# Virtual Environment
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
|
ENV/
|
||||||
|
env/
|
||||||
|
|
||||||
|
# IDE
|
||||||
|
.vscode/
|
||||||
|
.idea/
|
||||||
|
*.swp
|
||||||
|
*.swo
|
||||||
|
*~
|
||||||
|
|
||||||
|
# Logs
|
||||||
|
*.log
|
||||||
|
logs/
|
||||||
|
|
||||||
|
# OS
|
||||||
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
|
||||||
|
# Data files
|
||||||
|
data/
|
||||||
|
output/
|
||||||
|
logs/
|
||||||
|
|
||||||
|
# Configuration files (entire directory)
|
||||||
|
config/
|
||||||
|
|
||||||
|
# Neo4j (entire directory)
|
||||||
|
neo4j/
|
||||||
|
|
||||||
|
# Ontology files
|
||||||
|
ontology/
|
||||||
|
|
||||||
|
# Documentation
|
||||||
|
*.md
|
||||||
|
|
||||||
|
# Batch files
|
||||||
|
*.bat
|
||||||
|
|
||||||
|
# Jupyter
|
||||||
|
.ipynb_checkpoints/
|
||||||
|
*.ipynb
|
||||||
|
|
||||||
|
# Testing
|
||||||
|
.pytest_cache/
|
||||||
|
.coverage
|
||||||
|
htmlcov/
|
||||||
|
.tox/
|
||||||
@@ -0,0 +1,155 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
整合基础政务实体和深层语义实体
|
||||||
|
方案3:合并项目节点,完全整合
|
||||||
|
"""
|
||||||
|
|
||||||
|
import pandas as pd
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
def merge_kg_data():
|
||||||
|
"""整合知识图谱数据"""
|
||||||
|
|
||||||
|
output_dir = Path(__file__).parent / 'output'
|
||||||
|
|
||||||
|
# 读取4个文件
|
||||||
|
print("=== 读取原始文件 ===")
|
||||||
|
nodes_basic = pd.read_csv(output_dir / 'nodes.csv', encoding='utf-8-sig')
|
||||||
|
nodes_deep = pd.read_csv(output_dir / 'nodes_desc.csv', encoding='utf-8-sig')
|
||||||
|
rels_basic = pd.read_csv(output_dir / 'rels.csv', encoding='utf-8-sig')
|
||||||
|
rels_deep = pd.read_csv(output_dir / 'rels_desc.csv', encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"基础节点: {len(nodes_basic)}")
|
||||||
|
print(f"深层节点: {len(nodes_deep)}")
|
||||||
|
print(f"基础关系: {len(rels_basic)}")
|
||||||
|
print(f"深层关系: {len(rels_deep)}")
|
||||||
|
|
||||||
|
# 1. 合并ICH项目节点
|
||||||
|
print("\n=== 合并ICH项目节点 ===")
|
||||||
|
basic_projects = nodes_basic[nodes_basic['type'] == 'ICH_Project'].copy()
|
||||||
|
deep_projects = nodes_deep[nodes_deep['type'] == 'ICH_Project'].copy()
|
||||||
|
|
||||||
|
print(f"基础项目节点: {len(basic_projects)}")
|
||||||
|
print(f"深层项目节点: {len(deep_projects)}")
|
||||||
|
|
||||||
|
# 合并项目节点的属性
|
||||||
|
merged_projects = []
|
||||||
|
|
||||||
|
for pid in basic_projects['id']:
|
||||||
|
basic_row = basic_projects[basic_projects['id'] == pid].iloc[0]
|
||||||
|
|
||||||
|
# 查找深层项目节点
|
||||||
|
deep_row = deep_projects[deep_projects['id'] == pid]
|
||||||
|
|
||||||
|
if len(deep_row) > 0:
|
||||||
|
# 合并属性
|
||||||
|
deep_row = deep_row.iloc[0]
|
||||||
|
|
||||||
|
# 解析属性JSON
|
||||||
|
basic_props = json.loads(basic_row['properties']) if basic_row['properties'] else {}
|
||||||
|
deep_props = json.loads(deep_row['properties']) if deep_row['properties'] else {}
|
||||||
|
|
||||||
|
# 合并属性(深层属性优先)
|
||||||
|
merged_props = {**basic_props, **deep_props}
|
||||||
|
|
||||||
|
merged_projects.append({
|
||||||
|
'id': pid,
|
||||||
|
'label': basic_row['label'], # 保留基础标签
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': json.dumps(merged_props, ensure_ascii=False)
|
||||||
|
})
|
||||||
|
else:
|
||||||
|
# 只在基础中存在
|
||||||
|
merged_projects.append({
|
||||||
|
'id': pid,
|
||||||
|
'label': basic_row['label'],
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': basic_row['properties']
|
||||||
|
})
|
||||||
|
|
||||||
|
print(f"合并后项目节点: {len(merged_projects)}")
|
||||||
|
|
||||||
|
# 2. 合并其他节点(去除ICH_Project)
|
||||||
|
print("\n=== 合并其他节点 ===")
|
||||||
|
other_basic_nodes = nodes_basic[nodes_basic['type'] != 'ICH_Project']
|
||||||
|
other_deep_nodes = nodes_deep[nodes_deep['type'] != 'ICH_Project']
|
||||||
|
|
||||||
|
all_nodes = pd.concat([
|
||||||
|
pd.DataFrame(merged_projects),
|
||||||
|
other_basic_nodes,
|
||||||
|
other_deep_nodes
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
print(f"合并后总节点数: {len(all_nodes)}")
|
||||||
|
print(f" - ICH_Project: {len(merged_projects)}")
|
||||||
|
print(f" - 其他节点: {len(all_nodes) - len(merged_projects)}")
|
||||||
|
|
||||||
|
# 3. 合并关系
|
||||||
|
print("\n=== 合并关系 ===")
|
||||||
|
all_rels = pd.concat([rels_basic, rels_deep], ignore_index=True)
|
||||||
|
|
||||||
|
# 去重关系
|
||||||
|
all_rels = all_rels.drop_duplicates(subset=['source', 'target', 'type'], keep='first')
|
||||||
|
|
||||||
|
print(f"合并后关系数: {len(all_rels)}")
|
||||||
|
|
||||||
|
# 4. 统计信息
|
||||||
|
print("\n=== 整合后统计 ===")
|
||||||
|
print(f"总节点数: {len(all_nodes)}")
|
||||||
|
print(f"总关系数: {len(all_rels)}")
|
||||||
|
|
||||||
|
print("\n节点类型分布:")
|
||||||
|
for ntype, count in all_nodes.groupby('type').size().items():
|
||||||
|
print(f" {ntype}: {count}")
|
||||||
|
|
||||||
|
print(f"\n关系类型数量: {len(all_rels['type'].unique())}")
|
||||||
|
|
||||||
|
# 5. 保存整合后的文件
|
||||||
|
print("\n=== 保存整合文件 ===")
|
||||||
|
output_file_nodes = output_dir / 'kg_merged_nodes.csv'
|
||||||
|
output_file_rels = output_dir / 'kg_merged_rels.csv'
|
||||||
|
|
||||||
|
all_nodes.to_csv(output_file_nodes, index=False, encoding='utf-8-sig')
|
||||||
|
all_rels.to_csv(output_file_rels, index=False, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"节点已保存: {output_file_nodes}")
|
||||||
|
print(f"关系已保存: {output_file_rels}")
|
||||||
|
|
||||||
|
# 6. 验证
|
||||||
|
print("\n=== 验证 ===")
|
||||||
|
|
||||||
|
# 检查项目节点完整性
|
||||||
|
project_count = len(all_nodes[all_nodes['type'] == 'ICH_Project'])
|
||||||
|
print(f"ICH项目节点数: {project_count} (应该是268)")
|
||||||
|
|
||||||
|
# 检查关系完整性
|
||||||
|
rel_sources = set(all_rels['source'].unique())
|
||||||
|
node_ids = set(all_nodes['id'].unique())
|
||||||
|
|
||||||
|
missing_sources = rel_sources - node_ids
|
||||||
|
if missing_sources:
|
||||||
|
print(f"警告: {len(missing_sources)} 个关系的源节点不在节点文件中")
|
||||||
|
else:
|
||||||
|
print("所有关系的源节点都存在于节点文件中")
|
||||||
|
|
||||||
|
# 统计每个项目的关系数
|
||||||
|
project_rels = all_rels[all_rels['source'].str.startswith('ICH-')].groupby('source').size()
|
||||||
|
print(f"\n有关系的项目数: {len(project_rels)}/268")
|
||||||
|
|
||||||
|
# 展示几个示例项目的统计
|
||||||
|
print("\n示例项目关系统计:")
|
||||||
|
sample_projects = ['ICH-1', 'ICH-118', 'ICH-229']
|
||||||
|
for pid in sample_projects:
|
||||||
|
basic_rel_count = len(rels_basic[rels_basic['source'] == pid])
|
||||||
|
deep_rel_count = len(rels_deep[rels_deep['source'] == pid])
|
||||||
|
total_rel_count = len(all_rels[all_rels['source'] == pid])
|
||||||
|
print(f" {pid}: 基础{basic_rel_count} + 深层{deep_rel_count} = 总计{total_rel_count} 条关系")
|
||||||
|
|
||||||
|
print("\n完成!")
|
||||||
|
|
||||||
|
return all_nodes, all_rels
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
merge_kg_data()
|
||||||
@@ -0,0 +1,203 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
将重新抽取的结果合并到现有的节点和关系文件中
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import pandas as pd
|
||||||
|
from pathlib import Path
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 添加模块路径
|
||||||
|
sys.path.append(str(Path(__file__).parent / 'src'))
|
||||||
|
|
||||||
|
from data_processing.entity_normalizer import EntityNormalizer
|
||||||
|
from data_processing.relationship_builder import RelationshipBuilder
|
||||||
|
|
||||||
|
def merge_retry_results():
|
||||||
|
"""合并重新抽取的结果"""
|
||||||
|
|
||||||
|
# 文件路径
|
||||||
|
retry_file = Path(__file__).parent / 'output' / 'retry_projects.json'
|
||||||
|
nodes_file = Path(__file__).parent / 'output' / 'nodes_llm.csv'
|
||||||
|
rels_file = Path(__file__).parent / 'output' / 'rels_llm.csv'
|
||||||
|
ontology_file = Path(__file__).parent / 'config' / 'entity_ontology.yaml'
|
||||||
|
|
||||||
|
print("=== 加载数据 ===")
|
||||||
|
|
||||||
|
# 读取重新抽取的结果
|
||||||
|
with open(retry_file, 'r', encoding='utf-8') as f:
|
||||||
|
retry_data = json.load(f)
|
||||||
|
|
||||||
|
print(f"重新抽取的项目数: {len(retry_data)}")
|
||||||
|
|
||||||
|
# 读取现有的节点和关系
|
||||||
|
existing_nodes_df = pd.read_csv(nodes_file, encoding='utf-8-sig')
|
||||||
|
existing_rels_df = pd.read_csv(rels_file, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"现有节点数: {len(existing_nodes_df)}")
|
||||||
|
print(f"现有关系数: {len(existing_rels_df)}")
|
||||||
|
|
||||||
|
# 初始化规范化器和关系构建器
|
||||||
|
normalizer = EntityNormalizer(str(ontology_file))
|
||||||
|
builder = RelationshipBuilder(str(ontology_file))
|
||||||
|
|
||||||
|
# 规范化重新抽取的实体
|
||||||
|
print("\n=== 规范化实体 ===")
|
||||||
|
|
||||||
|
# 收集所有需要规范化的实体
|
||||||
|
all_entities_to_normalize = []
|
||||||
|
for item in retry_data:
|
||||||
|
result = item['extraction_result']
|
||||||
|
entities = result.get('entities', [])
|
||||||
|
# 添加项目ID以便跟踪
|
||||||
|
for entity in entities:
|
||||||
|
entity['_project_id'] = item['project_id']
|
||||||
|
all_entities_to_normalize.extend(entities)
|
||||||
|
|
||||||
|
# 批量规范化
|
||||||
|
extraction_results = []
|
||||||
|
for item in retry_data:
|
||||||
|
extraction_results.append(item['extraction_result'])
|
||||||
|
|
||||||
|
normalizer.normalize_batch(extraction_results)
|
||||||
|
|
||||||
|
# 生成新的实体节点
|
||||||
|
new_entity_nodes = normalizer.get_entity_nodes()
|
||||||
|
print(f"新生成实体节点数: {len(new_entity_nodes)}")
|
||||||
|
|
||||||
|
# 添加项目节点
|
||||||
|
new_project_nodes = []
|
||||||
|
for item in retry_data:
|
||||||
|
new_project_nodes.append({
|
||||||
|
'id': item['project_id'],
|
||||||
|
'label': item['project_name'],
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
print(f"新增项目节点数: {len(new_project_nodes)}")
|
||||||
|
|
||||||
|
# 构建关系
|
||||||
|
print("\n=== 构建关系 ===")
|
||||||
|
all_new_relationships = []
|
||||||
|
|
||||||
|
# 获取当前所有节点的ID到标签映射(用于查找目标实体ID)
|
||||||
|
all_nodes_for_lookup = pd.concat([
|
||||||
|
existing_nodes_df,
|
||||||
|
pd.DataFrame(new_entity_nodes)
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
# 创建名称->ID的映射字典
|
||||||
|
name_to_id = {}
|
||||||
|
for idx, row in all_nodes_for_lookup.iterrows():
|
||||||
|
if row['type'] != 'ICH_Project': # 只映射实体节点
|
||||||
|
name_to_id[(row['type'], row['label'])] = row['id']
|
||||||
|
|
||||||
|
for item in retry_data:
|
||||||
|
project_id = item['project_id']
|
||||||
|
result = item['extraction_result']
|
||||||
|
|
||||||
|
entities = result.get('entities', [])
|
||||||
|
relationships = result.get('relationships', [])
|
||||||
|
|
||||||
|
# 构建实体名称到ID的映射(仅限当前项目)
|
||||||
|
entity_name_to_id = {}
|
||||||
|
for entity in entities:
|
||||||
|
# 实体名称可能在name字段或attributes.name字段
|
||||||
|
entity_name = entity.get('name') or entity.get('attributes', {}).get('name', '')
|
||||||
|
entity_type = entity.get('type', '')
|
||||||
|
normalized_name = normalizer.normalize_text(entity_name)
|
||||||
|
|
||||||
|
# 在规范化器的映射中查找(键是normalized_text,不是元组)
|
||||||
|
entity_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
if not entity_id:
|
||||||
|
# 在所有节点中查找(可能是已存在的实体)
|
||||||
|
entity_id = name_to_id.get((entity_type, entity_name))
|
||||||
|
|
||||||
|
if entity_id:
|
||||||
|
entity_name_to_id[entity_name] = entity_id
|
||||||
|
|
||||||
|
# 手动构建关系
|
||||||
|
for rel in relationships:
|
||||||
|
# 关系中的字段名是source_entity和target_entity
|
||||||
|
source_name = rel.get('source_entity') or rel.get('source')
|
||||||
|
target_name = rel.get('target_entity') or rel.get('target')
|
||||||
|
rel_type = rel.get('type')
|
||||||
|
rel_props = rel.get('properties', {})
|
||||||
|
|
||||||
|
# 查找源实体ID
|
||||||
|
if source_name == project_id:
|
||||||
|
source_id = project_id
|
||||||
|
else:
|
||||||
|
# 先尝试在当前项目实体中查找
|
||||||
|
source_id = entity_name_to_id.get(source_name)
|
||||||
|
if not source_id:
|
||||||
|
# 在规范化器的映射中查找
|
||||||
|
normalized_name = normalizer.normalize_text(source_name)
|
||||||
|
source_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
|
||||||
|
# 查找目标实体ID
|
||||||
|
# 先尝试在当前项目实体中查找
|
||||||
|
target_id = entity_name_to_id.get(target_name)
|
||||||
|
if not target_id:
|
||||||
|
# 在规范化器的映射中查找
|
||||||
|
normalized_name = normalizer.normalize_text(target_name)
|
||||||
|
target_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
|
||||||
|
# 如果都找到了,添加关系
|
||||||
|
if source_id and target_id:
|
||||||
|
all_new_relationships.append({
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type,
|
||||||
|
'properties': json.dumps(rel_props, ensure_ascii=False) if rel_props else '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
print(f"项目 {project_id}: {len([r for r in all_new_relationships if r['source'] == project_id])} 条关系")
|
||||||
|
|
||||||
|
print(f"新增关系总数: {len(all_new_relationships)}")
|
||||||
|
|
||||||
|
# 合并节点
|
||||||
|
print("\n=== 合并节点 ===")
|
||||||
|
# 过滤掉已经存在的项目节点
|
||||||
|
existing_project_ids = set(existing_nodes_df[existing_nodes_df['type'] == 'ICH_Project']['id'].tolist())
|
||||||
|
new_project_nodes_filtered = [n for n in new_project_nodes if n['id'] not in existing_project_ids]
|
||||||
|
|
||||||
|
all_nodes = pd.concat([
|
||||||
|
existing_nodes_df,
|
||||||
|
pd.DataFrame(new_entity_nodes),
|
||||||
|
pd.DataFrame(new_project_nodes_filtered)
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
print(f"合并后节点数: {len(all_nodes)} (新增 {len(new_entity_nodes) + len(new_project_nodes_filtered)} 个)")
|
||||||
|
|
||||||
|
# 合并关系
|
||||||
|
print("\n=== 合并关系 ===")
|
||||||
|
all_rels = pd.concat([
|
||||||
|
existing_rels_df,
|
||||||
|
pd.DataFrame(all_new_relationships)
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
print(f"合并后关系数: {len(all_rels)} (新增 {len(all_new_relationships)} 条)")
|
||||||
|
|
||||||
|
# 保存结果
|
||||||
|
print("\n=== 保存结果 ===")
|
||||||
|
all_nodes.to_csv(nodes_file, index=False, encoding='utf-8-sig')
|
||||||
|
all_rels.to_csv(rels_file, index=False, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"节点已保存: {nodes_file}")
|
||||||
|
print(f"关系已保存: {rels_file}")
|
||||||
|
|
||||||
|
# 验证
|
||||||
|
print("\n=== 验证 ===")
|
||||||
|
for item in retry_data:
|
||||||
|
project_id = item['project_id']
|
||||||
|
rel_count = len(all_rels[all_rels['source'] == project_id])
|
||||||
|
print(f"{project_id}: {rel_count} 条关系")
|
||||||
|
|
||||||
|
print("\n完成!")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
merge_retry_results()
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
# 黑龙江省非物质文化遗产知识图谱构建 - 依赖包
|
||||||
|
|
||||||
|
# DeepSeek和LLM相关
|
||||||
|
langchain-deepseek>=0.1.0
|
||||||
|
langchain-core>=0.1.0
|
||||||
|
openai>=1.0.0 # 备用
|
||||||
|
|
||||||
|
# Neo4j数据库
|
||||||
|
neo4j>=5.15.0
|
||||||
|
py2neo>=2021.2.4
|
||||||
|
|
||||||
|
# 数据处理
|
||||||
|
pandas>=2.0.0
|
||||||
|
numpy>=1.24.0
|
||||||
|
openpyxl>=3.1.0 # Excel读取
|
||||||
|
|
||||||
|
# NLP和文本处理
|
||||||
|
jieba>=0.42.0
|
||||||
|
# 可选:深度学习框架
|
||||||
|
# torch>=2.0.0
|
||||||
|
# transformers>=4.30.0
|
||||||
|
|
||||||
|
# Web框架(用于后续开发)
|
||||||
|
# fastapi>=0.100.0
|
||||||
|
# uvicorn>=0.23.0
|
||||||
|
# flask>=3.0.0
|
||||||
|
|
||||||
|
# 可视化(用于后续开发)
|
||||||
|
# matplotlib>=3.7.0
|
||||||
|
# seaborn>=0.12.0
|
||||||
|
|
||||||
|
# 配置和日志
|
||||||
|
pyyaml>=6.0
|
||||||
|
python-dotenv>=1.0.0
|
||||||
|
|
||||||
|
# 其他工具
|
||||||
|
tqdm>=4.65.0
|
||||||
|
requests>=2.31.0
|
||||||
@@ -0,0 +1,92 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
重新抽取失败项目的脚本
|
||||||
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import pandas as pd
|
||||||
|
import yaml
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from datetime import datetime
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 添加模块路径
|
||||||
|
sys.path.append(str(Path(__file__).parent / 'src'))
|
||||||
|
|
||||||
|
from knowledge_extraction.deep_entity_extractor import DeepEntityExtractor
|
||||||
|
|
||||||
|
async def retry_failed_projects():
|
||||||
|
"""重新抽取失败的项目"""
|
||||||
|
|
||||||
|
# 加载配置
|
||||||
|
config_file = Path(__file__).parent / 'config' / 'deep_extraction_config.yaml'
|
||||||
|
with open(config_file, 'r', encoding='utf-8') as f:
|
||||||
|
config = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 读取原始数据
|
||||||
|
data_file = Path(__file__).parent.parent.parent / 'data' / '黑龙江国家级和省级非遗名单.xlsx'
|
||||||
|
df = pd.read_excel(data_file, engine='openpyxl')
|
||||||
|
|
||||||
|
# 找到失败的项目
|
||||||
|
# 第0列是序号,需要拼接成ICH-xxx格式
|
||||||
|
failed_data = df[df.iloc[:, 0].isin([118, 229])]
|
||||||
|
|
||||||
|
print(f"找到 {len(failed_data)} 个失败项目")
|
||||||
|
print("=" * 60)
|
||||||
|
|
||||||
|
# 初始化抽取器
|
||||||
|
extractor = DeepEntityExtractor(str(config_file))
|
||||||
|
|
||||||
|
results = []
|
||||||
|
|
||||||
|
for idx, row in failed_data.iterrows():
|
||||||
|
project_num = int(row.iloc[0]) # 序号(118, 229)
|
||||||
|
project_id = f'ICH-{project_num}' # 拼接成ICH-xxx格式
|
||||||
|
project_name = row.iloc[3] # 项目名称(第4列)
|
||||||
|
description = row.iloc[7] if len(row) > 7 else "" # 完整描述(第8列备注)
|
||||||
|
|
||||||
|
print(f"\n正在抽取: {project_id} - {project_name}")
|
||||||
|
print(f"描述长度: {len(description)} 字符")
|
||||||
|
|
||||||
|
# 准备输入数据
|
||||||
|
input_data = {
|
||||||
|
'project_id': project_id,
|
||||||
|
'project_name': project_name,
|
||||||
|
'description': description
|
||||||
|
}
|
||||||
|
|
||||||
|
# 抽取实体和关系
|
||||||
|
try:
|
||||||
|
result = await extractor.extract_from_remark(
|
||||||
|
project_id=project_id,
|
||||||
|
project_name=project_name,
|
||||||
|
remark_text=description
|
||||||
|
)
|
||||||
|
|
||||||
|
if result:
|
||||||
|
results.append({
|
||||||
|
'project_id': project_id,
|
||||||
|
'project_name': project_name,
|
||||||
|
'extraction_result': result
|
||||||
|
})
|
||||||
|
print(f"[OK] 抽取成功: {len(result.get('entities', []))} 个实体, {len(result.get('relationships', []))} 条关系")
|
||||||
|
else:
|
||||||
|
print(f"[FAIL] 抽取失败")
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
print(f"[ERROR] 抽取异常: {str(e)}")
|
||||||
|
|
||||||
|
# 保存结果
|
||||||
|
output_file = Path(__file__).parent / 'output' / 'retry_projects.json'
|
||||||
|
with open(output_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(results, f, ensure_ascii=False, indent=2)
|
||||||
|
|
||||||
|
print(f"\n结果已保存到: {output_file}")
|
||||||
|
print(f"成功: {len(results)}/{len(failed_data)}")
|
||||||
|
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
asyncio.run(retry_failed_projects())
|
||||||
@@ -0,0 +1,241 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
重新抽取失败项目的脚本
|
||||||
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import pandas as pd
|
||||||
|
import yaml
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from datetime import datetime
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 添加模块路径
|
||||||
|
sys.path.append(str(Path(__file__).parent / 'src'))
|
||||||
|
|
||||||
|
from knowledge_extraction.deep_entity_extractor import DeepEntityExtractor
|
||||||
|
from data_processing.entity_normalizer import EntityNormalizer
|
||||||
|
from data_processing.relationship_builder import RelationshipBuilder
|
||||||
|
|
||||||
|
async def retry_and_merge():
|
||||||
|
"""重新抽取失败项目并直接合并到现有文件"""
|
||||||
|
|
||||||
|
# 加载配置
|
||||||
|
config_file = Path(__file__).parent / 'config' / 'deep_extraction_config.yaml'
|
||||||
|
with open(config_file, 'r', encoding='utf-8') as f:
|
||||||
|
config = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 读取原始数据
|
||||||
|
data_file = Path(__file__).parent.parent.parent / 'data' / '黑龙江国家级和省级非遗名单.xlsx'
|
||||||
|
df = pd.read_excel(data_file, engine='openpyxl')
|
||||||
|
|
||||||
|
# 找到失败的项目
|
||||||
|
failed_data = df[df.iloc[:, 0].isin([118, 229])]
|
||||||
|
|
||||||
|
print(f"找到 {len(failed_data)} 个失败项目")
|
||||||
|
print("=" * 60)
|
||||||
|
|
||||||
|
# 初始化组件
|
||||||
|
extractor = DeepEntityExtractor(str(config_file))
|
||||||
|
ontology_file = Path(__file__).parent / 'config' / 'entity_ontology.yaml'
|
||||||
|
normalizer = EntityNormalizer(str(ontology_file))
|
||||||
|
builder = RelationshipBuilder(str(ontology_file))
|
||||||
|
|
||||||
|
# 读取现有的实体注册表(如果存在)
|
||||||
|
existing_nodes_file = Path(__file__).parent / 'output' / 'nodes_llm.csv'
|
||||||
|
existing_rels_file = Path(__file__).parent / 'output' / 'rels_llm.csv'
|
||||||
|
|
||||||
|
existing_nodes_df = pd.read_csv(existing_nodes_file, encoding='utf-8-sig')
|
||||||
|
existing_rels_df = pd.read_csv(existing_rels_file, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"现有节点数: {len(existing_nodes_df)}")
|
||||||
|
print(f"现有关系数: {len(existing_rels_df)}")
|
||||||
|
|
||||||
|
# 将现有实体加载到规范化器中
|
||||||
|
print("\n=== 加载现有实体到规范化器 ===")
|
||||||
|
for idx, row in existing_nodes_df.iterrows():
|
||||||
|
if row['type'] != 'ICH_Project':
|
||||||
|
# 将现有实体添加到规范化器的注册表
|
||||||
|
entity_type = row['type']
|
||||||
|
entity_text = row['label']
|
||||||
|
entity_id = row['id']
|
||||||
|
|
||||||
|
# 规范化文本
|
||||||
|
normalized_text = normalizer.normalize_text(entity_text)
|
||||||
|
|
||||||
|
# 添加到映射表
|
||||||
|
if normalized_text not in normalizer.text_to_id_map:
|
||||||
|
normalizer.text_to_id_map[normalized_text] = entity_id
|
||||||
|
normalizer.entity_registry[entity_id] = {
|
||||||
|
'type': entity_type,
|
||||||
|
'text': entity_text,
|
||||||
|
'normalized_text': normalized_text,
|
||||||
|
'canonical_name': normalized_text # 添加这个字段
|
||||||
|
}
|
||||||
|
|
||||||
|
print(f"已加载 {len(normalizer.entity_registry)} 个现有实体")
|
||||||
|
|
||||||
|
# 重新抽取失败的项目
|
||||||
|
extraction_results = []
|
||||||
|
|
||||||
|
for idx, row in failed_data.iterrows():
|
||||||
|
project_num = int(row.iloc[0])
|
||||||
|
project_id = f'ICH-{project_num}'
|
||||||
|
project_name = row.iloc[3]
|
||||||
|
description = row.iloc[7] if len(row) > 7 else ""
|
||||||
|
|
||||||
|
print(f"\n正在抽取: {project_id} - {project_name}")
|
||||||
|
print(f"描述长度: {len(description)} 字符")
|
||||||
|
|
||||||
|
try:
|
||||||
|
result = await extractor.extract_from_remark(
|
||||||
|
project_id=project_id,
|
||||||
|
project_name=project_name,
|
||||||
|
remark_text=description
|
||||||
|
)
|
||||||
|
|
||||||
|
if result:
|
||||||
|
extraction_results.append({
|
||||||
|
'project_id': project_id,
|
||||||
|
'project_name': project_name,
|
||||||
|
'extraction_result': result
|
||||||
|
})
|
||||||
|
print(f"[OK] 抽取成功: {len(result.get('entities', []))} 个实体, {len(result.get('relationships', []))} 条关系")
|
||||||
|
else:
|
||||||
|
print(f"[FAIL] 抽取失败")
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
print(f"[ERROR] 抽取异常: {str(e)}")
|
||||||
|
|
||||||
|
if not extraction_results:
|
||||||
|
print("\n没有成功抽取的项目")
|
||||||
|
return
|
||||||
|
|
||||||
|
# 规范化新抽取的实体(会自动去重)
|
||||||
|
print("\n=== 规范化新抽取的实体 ===")
|
||||||
|
for item in extraction_results:
|
||||||
|
result = item['extraction_result']
|
||||||
|
entities = result.get('entities', [])
|
||||||
|
|
||||||
|
for entity in entities:
|
||||||
|
# 实体名称可能在name字段或attributes.name字段
|
||||||
|
entity_name = entity.get('name') or entity.get('attributes', {}).get('name', '')
|
||||||
|
entity_type = entity.get('type', '')
|
||||||
|
|
||||||
|
# 规范化实体(会自动去重)
|
||||||
|
normalizer.normalize_entity(entity, similarity_threshold=0.85)
|
||||||
|
|
||||||
|
# 获取所有实体节点(包括新增的)
|
||||||
|
all_entity_nodes = normalizer.get_entity_nodes()
|
||||||
|
print(f"规范化后实体节点数: {len(all_entity_nodes)}")
|
||||||
|
|
||||||
|
# 构建关系
|
||||||
|
print("\n=== 构建关系 ===")
|
||||||
|
all_new_relationships = []
|
||||||
|
|
||||||
|
for item in extraction_results:
|
||||||
|
project_id = item['project_id']
|
||||||
|
result = item['extraction_result']
|
||||||
|
|
||||||
|
entities = result.get('entities', [])
|
||||||
|
relationships = result.get('relationships', [])
|
||||||
|
|
||||||
|
# 构建实体名称到ID的映射
|
||||||
|
entity_name_to_id = {}
|
||||||
|
for entity in entities:
|
||||||
|
entity_name = entity.get('name') or entity.get('attributes', {}).get('name', '')
|
||||||
|
entity_type = entity.get('type', '')
|
||||||
|
normalized_name = normalizer.normalize_text(entity_name)
|
||||||
|
entity_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
if entity_id:
|
||||||
|
entity_name_to_id[entity_name] = entity_id
|
||||||
|
|
||||||
|
# 手动构建关系
|
||||||
|
for rel in relationships:
|
||||||
|
source_name = rel.get('source_entity') or rel.get('source')
|
||||||
|
target_name = rel.get('target_entity') or rel.get('target')
|
||||||
|
rel_type = rel.get('type')
|
||||||
|
rel_props = rel.get('properties', {})
|
||||||
|
|
||||||
|
# 查找源实体ID
|
||||||
|
if source_name == project_id:
|
||||||
|
source_id = project_id
|
||||||
|
else:
|
||||||
|
normalized_name = normalizer.normalize_text(source_name)
|
||||||
|
source_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
|
||||||
|
# 查找目标实体ID
|
||||||
|
normalized_name = normalizer.normalize_text(target_name)
|
||||||
|
target_id = normalizer.text_to_id_map.get(normalized_name)
|
||||||
|
|
||||||
|
# 如果都找到了,添加关系
|
||||||
|
if source_id and target_id:
|
||||||
|
all_new_relationships.append({
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type,
|
||||||
|
'properties': json.dumps(rel_props, ensure_ascii=False) if rel_props else '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
print(f"项目 {project_id}: {len([r for r in all_new_relationships if r['source'] == project_id])} 条关系")
|
||||||
|
|
||||||
|
# 合并节点和关系
|
||||||
|
print("\n=== 合并数据 ===")
|
||||||
|
|
||||||
|
# 项目节点:只添加缺失的项目
|
||||||
|
existing_project_ids = set(existing_nodes_df[existing_nodes_df['type'] == 'ICH_Project']['id'].tolist())
|
||||||
|
new_project_nodes = [
|
||||||
|
{'id': item['project_id'], 'label': item['project_name'], 'type': 'ICH_Project', 'properties': '{}'}
|
||||||
|
for item in extraction_results
|
||||||
|
if item['project_id'] not in existing_project_ids
|
||||||
|
]
|
||||||
|
|
||||||
|
# 合并所有节点
|
||||||
|
all_nodes_df = pd.concat([
|
||||||
|
existing_nodes_df[existing_nodes_df['type'] != 'ICH_Project'], # 现有实体节点
|
||||||
|
pd.DataFrame(all_entity_nodes), # 所有实体节点(包括新增和去重后的现有)
|
||||||
|
pd.DataFrame(new_project_nodes), # 新增项目节点
|
||||||
|
existing_nodes_df[existing_nodes_df['type'] == 'ICH_Project'] # 现有项目节点
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
# 去重节点(按ID)
|
||||||
|
all_nodes_df = all_nodes_df.drop_duplicates(subset=['id'], keep='first')
|
||||||
|
|
||||||
|
# 合并关系
|
||||||
|
all_rels_df = pd.concat([
|
||||||
|
existing_rels_df,
|
||||||
|
pd.DataFrame(all_new_relationships)
|
||||||
|
], ignore_index=True)
|
||||||
|
|
||||||
|
# 去重关系
|
||||||
|
all_rels_df = all_rels_df.drop_duplicates(subset=['source', 'target', 'type'], keep='first')
|
||||||
|
|
||||||
|
print(f"合并后节点数: {len(all_nodes_df)} (新增 {len(all_entity_nodes) - len(existing_nodes_df[existing_nodes_df['type'] != 'ICH_Project'])} 个实体)")
|
||||||
|
print(f"合并后关系数: {len(all_rels_df)} (新增 {len(all_new_relationships)} 条)")
|
||||||
|
|
||||||
|
# 保存结果
|
||||||
|
print("\n=== 保存结果 ===")
|
||||||
|
all_nodes_df.to_csv(existing_nodes_file, index=False, encoding='utf-8-sig')
|
||||||
|
all_rels_df.to_csv(existing_rels_file, index=False, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f"节点已保存: {existing_nodes_file}")
|
||||||
|
print(f"关系已保存: {existing_rels_file}")
|
||||||
|
|
||||||
|
# 验证
|
||||||
|
print("\n=== 验证 ===")
|
||||||
|
for item in extraction_results:
|
||||||
|
project_id = item['project_id']
|
||||||
|
rel_count = len(all_rels_df[all_rels_df['source'] == project_id])
|
||||||
|
print(f"{project_id}: {rel_count} 条关系")
|
||||||
|
|
||||||
|
print("\n完成!")
|
||||||
|
|
||||||
|
# 保存抽取结果以供检查
|
||||||
|
output_file = Path(__file__).parent / 'output' / 'retry_projects.json'
|
||||||
|
with open(output_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(extraction_results, f, ensure_ascii=False, indent=2)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
asyncio.run(retry_and_merge())
|
||||||
@@ -0,0 +1,375 @@
|
|||||||
|
"""
|
||||||
|
数据预处理脚本 - 读取黑龙江非遗Excel数据
|
||||||
|
"""
|
||||||
|
|
||||||
|
import pandas as pd
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Dict, List, Any
|
||||||
|
import logging
|
||||||
|
|
||||||
|
class ExcelDataReader:
|
||||||
|
"""Excel数据读取器"""
|
||||||
|
|
||||||
|
def __init__(self, excel_path: str):
|
||||||
|
"""
|
||||||
|
初始化数据读取器
|
||||||
|
|
||||||
|
Args:
|
||||||
|
excel_path: Excel文件路径
|
||||||
|
"""
|
||||||
|
self.excel_path = Path(excel_path)
|
||||||
|
self.data = None
|
||||||
|
self.logger = self._setup_logger()
|
||||||
|
|
||||||
|
def _setup_logger(self):
|
||||||
|
"""设置日志"""
|
||||||
|
logging.basicConfig(
|
||||||
|
level=logging.INFO,
|
||||||
|
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
|
||||||
|
)
|
||||||
|
return logging.getLogger(__name__)
|
||||||
|
|
||||||
|
def read_excel(self) -> pd.DataFrame:
|
||||||
|
"""
|
||||||
|
读取Excel文件
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
DataFrame: 数据框
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
self.logger.info(f"开始读取Excel文件: {self.excel_path}")
|
||||||
|
|
||||||
|
# 读取Excel文件
|
||||||
|
self.data = pd.read_excel(self.excel_path)
|
||||||
|
|
||||||
|
self.logger.info(f"成功读取 {len(self.data)} 行数据")
|
||||||
|
self.logger.info(f"列名: {list(self.data.columns)}")
|
||||||
|
|
||||||
|
# 显示前5行
|
||||||
|
self.logger.info("\n前5行数据:")
|
||||||
|
self.logger.info(self.data.head())
|
||||||
|
|
||||||
|
return self.data
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"读取Excel文件失败: {str(e)}")
|
||||||
|
raise
|
||||||
|
|
||||||
|
def analyze_data(self) -> Dict[str, Any]:
|
||||||
|
"""
|
||||||
|
分析数据
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Dict: 分析结果
|
||||||
|
"""
|
||||||
|
if self.data is None:
|
||||||
|
raise ValueError("请先读取Excel文件")
|
||||||
|
|
||||||
|
analysis = {
|
||||||
|
"total_records": len(self.data),
|
||||||
|
"columns": list(self.data.columns),
|
||||||
|
"column_types": {col: str(dtype) for col, dtype in self.data.dtypes.items()},
|
||||||
|
"missing_values": self.data.isnull().sum().to_dict(),
|
||||||
|
"statistics": {}
|
||||||
|
}
|
||||||
|
|
||||||
|
# 分析类别分布
|
||||||
|
if '类别' in self.data.columns:
|
||||||
|
category_counts = self.data['类别'].value_counts()
|
||||||
|
analysis['category_distribution'] = category_counts.to_dict()
|
||||||
|
|
||||||
|
# 分析级别分布
|
||||||
|
if '项目级别' in self.data.columns:
|
||||||
|
level_counts = self.data['项目级别'].value_counts()
|
||||||
|
analysis['level_distribution'] = level_counts.to_dict()
|
||||||
|
|
||||||
|
# 分析地域分布
|
||||||
|
if '项目申报单位/地区' in self.data.columns:
|
||||||
|
location_counts = self.data['项目申报单位/地区'].value_counts()
|
||||||
|
analysis['location_distribution'] = location_counts.head(20).to_dict()
|
||||||
|
|
||||||
|
# 传承人覆盖率
|
||||||
|
if '代表性传承人' in self.data.columns:
|
||||||
|
has_inheritor = self.data['代表性传承人'].notna().sum()
|
||||||
|
analysis['inheritor_coverage'] = {
|
||||||
|
"total": len(self.data),
|
||||||
|
"has_inheritor": int(has_inheritor),
|
||||||
|
"coverage_rate": float(has_inheritor / len(self.data) * 100)
|
||||||
|
}
|
||||||
|
|
||||||
|
return analysis
|
||||||
|
|
||||||
|
def clean_data(self) -> pd.DataFrame:
|
||||||
|
"""
|
||||||
|
清洗数据
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
DataFrame: 清洗后的数据
|
||||||
|
"""
|
||||||
|
if self.data is None:
|
||||||
|
raise ValueError("请先读取Excel文件")
|
||||||
|
|
||||||
|
self.logger.info("开始清洗数据")
|
||||||
|
|
||||||
|
# 去除空行
|
||||||
|
original_len = len(self.data)
|
||||||
|
self.data = self.data.dropna(how='all')
|
||||||
|
self.logger.info(f"去除空行: {original_len} -> {len(self.data)}")
|
||||||
|
|
||||||
|
# 填充缺失值
|
||||||
|
for col in self.data.columns:
|
||||||
|
if self.data[col].dtype == 'object':
|
||||||
|
self.data[col] = self.data[col].fillna('')
|
||||||
|
else:
|
||||||
|
self.data[col] = self.data[col].fillna(0)
|
||||||
|
|
||||||
|
# 去重(基于项目名称、项目批次、项目保护单位、代表性传承人四个字段)
|
||||||
|
dedup_cols = ['项目名称', '项目批次', '项目保护单位', '代表性传承人']
|
||||||
|
existing_cols = [col for col in dedup_cols if col in self.data.columns]
|
||||||
|
|
||||||
|
if existing_cols:
|
||||||
|
before_dedup = len(self.data)
|
||||||
|
self.data = self.data.drop_duplicates(subset=existing_cols, keep='first')
|
||||||
|
duplicate_count = before_dedup - len(self.data)
|
||||||
|
dedup_rate = (duplicate_count / before_dedup * 100) if before_dedup > 0 else 0
|
||||||
|
self.logger.info(f"去重(基于{len(existing_cols)}个字段: {', '.join(existing_cols)}): {before_dedup} -> {len(self.data)} (删除{duplicate_count}条,去重率{dedup_rate:.1f}%)")
|
||||||
|
|
||||||
|
return self.data
|
||||||
|
|
||||||
|
def convert_to_kg_format(self) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
转换为知识图谱格式
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List[Dict]: 知识图谱节点列表
|
||||||
|
"""
|
||||||
|
if self.data is None:
|
||||||
|
raise ValueError("请先读取Excel文件")
|
||||||
|
|
||||||
|
self.logger.info("转换为知识图谱格式")
|
||||||
|
|
||||||
|
kg_nodes = []
|
||||||
|
|
||||||
|
for idx, row in self.data.iterrows():
|
||||||
|
try:
|
||||||
|
# 创建非遗项目节点
|
||||||
|
project_node = {
|
||||||
|
"id": f"ICH-{idx:04d}",
|
||||||
|
"type": "ICH_Project",
|
||||||
|
"properties": {
|
||||||
|
"project_id": f"ICH-{idx:04d}",
|
||||||
|
"name": str(row.get('项目名称', '')),
|
||||||
|
"level": str(row.get('项目级别', '')),
|
||||||
|
"batch": str(row.get('批次号', '')),
|
||||||
|
"category": str(row.get('类别', '')),
|
||||||
|
"declaration_area": str(row.get('项目申报单位/地区', '')),
|
||||||
|
"description": str(row.get('备注', '')),
|
||||||
|
"source_row": idx + 2 # Excel行号(从2开始,第一行是表头)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
kg_nodes.append(project_node)
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.warning(f"转换第{idx}行数据失败: {str(e)}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
self.logger.info(f"成功转换 {len(kg_nodes)} 个节点")
|
||||||
|
return kg_nodes
|
||||||
|
|
||||||
|
def extract_inheritors(self) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
提取传承人信息
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List[Dict]: 传承人节点列表
|
||||||
|
"""
|
||||||
|
if self.data is None:
|
||||||
|
raise ValueError("请先读取Excel文件")
|
||||||
|
|
||||||
|
self.logger.info("提取传承人信息")
|
||||||
|
|
||||||
|
inheritors = []
|
||||||
|
inheritor_id = 0
|
||||||
|
|
||||||
|
for idx, row in self.data.iterrows():
|
||||||
|
inheritor_names = row.get('代表性传承人', '')
|
||||||
|
if pd.isna(inheritor_names) or not str(inheritor_names).strip():
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 处理多个传承人(用顿号分隔)
|
||||||
|
names = str(inheritor_names).replace('、', ',').replace(',', ',').split(',')
|
||||||
|
|
||||||
|
for name in names:
|
||||||
|
name = name.strip()
|
||||||
|
if not name:
|
||||||
|
continue
|
||||||
|
|
||||||
|
inheritor_id += 1
|
||||||
|
inheritor_node = {
|
||||||
|
"id": f"INH-{inheritor_id:04d}",
|
||||||
|
"type": "Inheritor",
|
||||||
|
"properties": {
|
||||||
|
"inheritor_id": f"INH-{inheritor_id:04d}",
|
||||||
|
"name": name,
|
||||||
|
"project_name": str(row.get('项目名称', '')),
|
||||||
|
"source_row": idx + 2
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
inheritors.append(inheritor_node)
|
||||||
|
|
||||||
|
self.logger.info(f"提取了 {len(inheritors)} 个传承人")
|
||||||
|
return inheritors
|
||||||
|
|
||||||
|
def extract_relations(self) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
提取关系
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List[Dict]: 关系列表
|
||||||
|
"""
|
||||||
|
if self.data is None:
|
||||||
|
raise ValueError("请先读取Excel文件")
|
||||||
|
|
||||||
|
self.logger.info("提取关系")
|
||||||
|
|
||||||
|
relations = []
|
||||||
|
relation_id = 0
|
||||||
|
|
||||||
|
for idx, row in self.data.iterrows():
|
||||||
|
project_id = f"ICH-{idx:04d}"
|
||||||
|
|
||||||
|
# 项目-类别关系
|
||||||
|
category = row.get('类别', '')
|
||||||
|
if pd.notna(category) and str(category).strip():
|
||||||
|
relation_id += 1
|
||||||
|
relations.append({
|
||||||
|
"id": f"REL-{relation_id:04d}",
|
||||||
|
"type": "belongs_to",
|
||||||
|
"from": project_id,
|
||||||
|
"to": f"CAT-{str(category)}",
|
||||||
|
"properties": {}
|
||||||
|
})
|
||||||
|
|
||||||
|
# 项目-传承人关系
|
||||||
|
inheritor_names = row.get('代表性传承人', '')
|
||||||
|
if pd.notna(inheritor_names) and str(inheritor_names).strip():
|
||||||
|
names = str(inheritor_names).replace('、', ',').replace(',', ',').split(',')
|
||||||
|
for name in names:
|
||||||
|
name = name.strip()
|
||||||
|
if name:
|
||||||
|
relation_id += 1
|
||||||
|
relations.append({
|
||||||
|
"id": f"REL-{relation_id:04d}",
|
||||||
|
"type": "has_inheritor",
|
||||||
|
"from": project_id,
|
||||||
|
"to": f"INH-{name}", # 简化处理
|
||||||
|
"properties": {}
|
||||||
|
})
|
||||||
|
|
||||||
|
self.logger.info(f"提取了 {len(relations)} 个关系")
|
||||||
|
return relations
|
||||||
|
|
||||||
|
def save_analysis_report(self, output_path: str):
|
||||||
|
"""
|
||||||
|
保存分析报告
|
||||||
|
|
||||||
|
Args:
|
||||||
|
output_path: 输出路径
|
||||||
|
"""
|
||||||
|
analysis = self.analyze_data()
|
||||||
|
|
||||||
|
with open(output_path, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(analysis, f, indent=2, ensure_ascii=False)
|
||||||
|
|
||||||
|
self.logger.info(f"分析报告已保存到: {output_path}")
|
||||||
|
|
||||||
|
def save_kg_data(self, nodes: List[Dict], relations: List[Dict], output_dir: str):
|
||||||
|
"""
|
||||||
|
保存知识图谱数据
|
||||||
|
|
||||||
|
Args:
|
||||||
|
nodes: 节点列表
|
||||||
|
relations: 关系列表
|
||||||
|
output_dir: 输出目录
|
||||||
|
"""
|
||||||
|
output_path = Path(output_dir)
|
||||||
|
output_path.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# 保存节点
|
||||||
|
nodes_file = output_path / "nodes.json"
|
||||||
|
with open(nodes_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(nodes, f, indent=2, ensure_ascii=False)
|
||||||
|
self.logger.info(f"节点已保存到: {nodes_file}")
|
||||||
|
|
||||||
|
# 保存关系
|
||||||
|
relations_file = output_path / "relations.json"
|
||||||
|
with open(relations_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(relations, f, indent=2, ensure_ascii=False)
|
||||||
|
self.logger.info(f"关系已保存到: {relations_file}")
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
"""主函数"""
|
||||||
|
# 配置
|
||||||
|
excel_path = r"E:\Project\2026_KG_ICH\data\黑龙江国家级和省级非遗名单.xlsx"
|
||||||
|
output_dir = r"E:\Project\2026_KG_ICH\dofile\kg_project\output"
|
||||||
|
|
||||||
|
# 创建读取器
|
||||||
|
reader = ExcelDataReader(excel_path)
|
||||||
|
|
||||||
|
# 读取数据
|
||||||
|
reader.read_excel()
|
||||||
|
|
||||||
|
# 分析数据
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("数据分析")
|
||||||
|
print("="*60)
|
||||||
|
analysis = reader.analyze_data()
|
||||||
|
print(f"\n总记录数: {analysis['total_records']}")
|
||||||
|
print(f"\n类别分布:")
|
||||||
|
for cat, count in analysis.get('category_distribution', {}).items():
|
||||||
|
print(f" {cat}: {count}")
|
||||||
|
|
||||||
|
print(f"\n传承人覆盖率: {analysis.get('inheritor_coverage', {}).get('coverage_rate', 0):.1f}%")
|
||||||
|
|
||||||
|
# 保存分析报告
|
||||||
|
Path(output_dir).mkdir(parents=True, exist_ok=True)
|
||||||
|
reader.save_analysis_report(f"{output_dir}/data_analysis.json")
|
||||||
|
|
||||||
|
# 清洗数据
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("数据清洗")
|
||||||
|
print("="*60)
|
||||||
|
reader.clean_data()
|
||||||
|
|
||||||
|
# 转换为知识图谱格式
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("转换为知识图谱格式")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
nodes = reader.convert_to_kg_format()
|
||||||
|
inheritors = reader.extract_inheritors()
|
||||||
|
relations = reader.extract_relations()
|
||||||
|
|
||||||
|
# 合并所有节点
|
||||||
|
all_nodes = nodes + inheritors
|
||||||
|
|
||||||
|
print(f"\n节点统计:")
|
||||||
|
print(f" 非遗项目节点: {len(nodes)}")
|
||||||
|
print(f" 传承人节点: {len(inheritors)}")
|
||||||
|
print(f" 总节点数: {len(all_nodes)}")
|
||||||
|
print(f" 关系数: {len(relations)}")
|
||||||
|
|
||||||
|
# 保存知识图谱数据
|
||||||
|
reader.save_kg_data(all_nodes, relations, output_dir)
|
||||||
|
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("处理完成!")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -0,0 +1,438 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
从Excel提取数据生成知识图谱CSV文件
|
||||||
|
保持原始数据表述不变
|
||||||
|
"""
|
||||||
|
|
||||||
|
import pandas as pd
|
||||||
|
import re
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
def extract_batches(df):
|
||||||
|
"""提取所有唯一的批次(保持原始表述)"""
|
||||||
|
batches = {}
|
||||||
|
batch_counter = {}
|
||||||
|
|
||||||
|
for batch_str in df['项目批次'].dropna().unique():
|
||||||
|
if not batch_str or str(batch_str).strip() == '':
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 解析批次字符串(可能包含多个批次)
|
||||||
|
batch_items = re.split(r'[,、,]', str(batch_str))
|
||||||
|
|
||||||
|
for item in batch_items:
|
||||||
|
item = item.strip()
|
||||||
|
if not item:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 生成批次ID(使用原始表述的哈希)
|
||||||
|
if item not in batch_counter:
|
||||||
|
batch_counter[item] = 1
|
||||||
|
else:
|
||||||
|
batch_counter[item] += 1
|
||||||
|
|
||||||
|
batch_id = f"BATCH-{abs(hash(item)) % 100000:05d}"
|
||||||
|
|
||||||
|
if batch_id not in batches:
|
||||||
|
batches[batch_id] = {
|
||||||
|
'name': item, # 保持原始表述
|
||||||
|
'original_string': item
|
||||||
|
}
|
||||||
|
|
||||||
|
return batches
|
||||||
|
|
||||||
|
|
||||||
|
def parse_inheritor_field(inheritor_str):
|
||||||
|
"""解析传承人字段,保持原始表述(如"吴明新(国)")"""
|
||||||
|
if not inheritor_str or str(inheritor_str).strip() in ['无', '']:
|
||||||
|
return []
|
||||||
|
|
||||||
|
# 按顿号、逗号分割,保持原始表述
|
||||||
|
inheritors = re.split(r'[、,,\n]', str(inheritor_str))
|
||||||
|
inheritors = [inh.strip() for inh in inheritors if inh.strip() and inh.strip() != '无']
|
||||||
|
return inheritors
|
||||||
|
|
||||||
|
|
||||||
|
def extract_inheritors(df):
|
||||||
|
"""提取所有唯一传承人(保持原始表述)"""
|
||||||
|
inheritors = {}
|
||||||
|
inheritor_counter = {}
|
||||||
|
|
||||||
|
for idx, row in df.iterrows():
|
||||||
|
inheritor_str = row.get('代表性传承人', '')
|
||||||
|
if not inheritor_str or str(inheritor_str).strip() in ['无', '']:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 解析传承人列表
|
||||||
|
inheritor_names = parse_inheritor_field(inheritor_str)
|
||||||
|
|
||||||
|
for name in inheritor_names:
|
||||||
|
# 使用原始名称(包括括号)作为key
|
||||||
|
if name not in inheritor_counter:
|
||||||
|
inheritor_counter[name] = 1
|
||||||
|
else:
|
||||||
|
inheritor_counter[name] += 1
|
||||||
|
|
||||||
|
# 生成传承人ID
|
||||||
|
inheritor_id = f"INH-{abs(hash(name)) % 100000:05d}-{inheritor_counter[name]}"
|
||||||
|
|
||||||
|
if inheritor_id not in inheritors:
|
||||||
|
inheritors[inheritor_id] = {
|
||||||
|
'name': name # 保持原始表述,如"吴明新(国)"
|
||||||
|
}
|
||||||
|
|
||||||
|
return inheritors
|
||||||
|
|
||||||
|
|
||||||
|
def extract_institutions(df):
|
||||||
|
"""提取所有唯一保护机构(保持原始表述)"""
|
||||||
|
institutions = {}
|
||||||
|
inst_counter = {}
|
||||||
|
|
||||||
|
for idx, row in df.iterrows():
|
||||||
|
inst_str = row.get('项目保护单位', '')
|
||||||
|
if not inst_str or str(inst_str).strip() in ['', '无']:
|
||||||
|
continue
|
||||||
|
|
||||||
|
inst_name = str(inst_str).strip()
|
||||||
|
|
||||||
|
# 使用原始机构名
|
||||||
|
if inst_name not in inst_counter:
|
||||||
|
inst_counter[inst_name] = 1
|
||||||
|
else:
|
||||||
|
inst_counter[inst_name] += 1
|
||||||
|
|
||||||
|
# 生成机构ID
|
||||||
|
inst_id = f"INST-{abs(hash(inst_name)) % 100000:05d}-{inst_counter[inst_name]}"
|
||||||
|
|
||||||
|
if inst_id not in institutions:
|
||||||
|
institutions[inst_id] = {
|
||||||
|
'name': inst_name # 保持原始表述
|
||||||
|
}
|
||||||
|
|
||||||
|
return institutions
|
||||||
|
|
||||||
|
|
||||||
|
def parse_batch_field(batch_str, batches_dict):
|
||||||
|
"""解析批次字段,返回批次ID列表"""
|
||||||
|
if not batch_str or str(batch_str).strip() == '':
|
||||||
|
return []
|
||||||
|
|
||||||
|
batches = re.split(r'[,、,]', str(batch_str))
|
||||||
|
batch_ids = []
|
||||||
|
|
||||||
|
for batch in batches:
|
||||||
|
batch = batch.strip()
|
||||||
|
if batch:
|
||||||
|
# 查找对应的批次ID
|
||||||
|
for batch_id, batch_info in batches_dict.items():
|
||||||
|
if batch_info['name'] == batch:
|
||||||
|
batch_ids.append(batch_id)
|
||||||
|
break
|
||||||
|
|
||||||
|
return batch_ids
|
||||||
|
|
||||||
|
|
||||||
|
def get_institution_id(institution_name, institutions_dict):
|
||||||
|
"""获取机构ID"""
|
||||||
|
if not institution_name or str(institution_name).strip() in ['', '无']:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 从已提取的机构中查找
|
||||||
|
for inst_id, inst_info in institutions_dict.items():
|
||||||
|
if inst_info['name'] == str(institution_name).strip():
|
||||||
|
return inst_id
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def get_inheritor_id(inheritor_name, inheritors_dict):
|
||||||
|
"""获取传承人ID"""
|
||||||
|
if not inheritor_name or not inheritor_name.strip():
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 从已提取的传承人中查找
|
||||||
|
for inheritor_id, inheritor_info in inheritors_dict.items():
|
||||||
|
if inheritor_info['name'] == inheritor_name:
|
||||||
|
return inheritor_id
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def clean_data(df):
|
||||||
|
"""数据清洗"""
|
||||||
|
# 去除完全空白的行
|
||||||
|
df = df.dropna(how='all')
|
||||||
|
|
||||||
|
# 填充缺失值为空字符串
|
||||||
|
for col in df.columns:
|
||||||
|
if df[col].dtype == 'object':
|
||||||
|
df[col] = df[col].fillna('')
|
||||||
|
|
||||||
|
return df
|
||||||
|
|
||||||
|
|
||||||
|
def build_nodes(df):
|
||||||
|
"""构建节点数据(保持原始表述)"""
|
||||||
|
nodes_list = []
|
||||||
|
|
||||||
|
print("正在构建节点...")
|
||||||
|
|
||||||
|
# 1. ICH_Project节点
|
||||||
|
print(" - 构建ICH_Project节点...")
|
||||||
|
for idx, row in df.iterrows():
|
||||||
|
seq_num = row.get('总序号', idx + 1)
|
||||||
|
project_name = row.get('项目名称', '')
|
||||||
|
|
||||||
|
# 确定级别(从批次字段推断)
|
||||||
|
batch_str = str(row.get('项目批次', ''))
|
||||||
|
level = '国家级' if '国家级' in batch_str else ('省级' if '省级' in batch_str else '')
|
||||||
|
|
||||||
|
nodes_list.append({
|
||||||
|
'id': f"ICH-{int(seq_num)}",
|
||||||
|
'label': str(project_name),
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': json.dumps({
|
||||||
|
'level': level,
|
||||||
|
'category': row.get('类别', ''), # 保持原始表述,如"曲艺类"
|
||||||
|
'batch': batch_str,
|
||||||
|
'protection_unit': row.get('项目保护单位', '')
|
||||||
|
}, ensure_ascii=False)
|
||||||
|
})
|
||||||
|
|
||||||
|
# 2. Category节点(从数据中动态提取,保持原始表述)
|
||||||
|
print(" - 构建Category节点...")
|
||||||
|
unique_categories = df['类别'].dropna().unique()
|
||||||
|
|
||||||
|
for cat_name in unique_categories:
|
||||||
|
# 使用类别名称作为ID(使用哈希避免特殊字符)
|
||||||
|
cat_id = f"CAT-{abs(hash(cat_name)) % 100000:05d}"
|
||||||
|
|
||||||
|
nodes_list.append({
|
||||||
|
'id': cat_id,
|
||||||
|
'label': cat_name, # 保持原始表述,如"曲艺类"
|
||||||
|
'type': 'Category',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
# 3. Batch节点(动态识别,保持原始表述)
|
||||||
|
print(" - 构建Batch节点...")
|
||||||
|
batches = extract_batches(df)
|
||||||
|
for batch_id, batch_info in batches.items():
|
||||||
|
nodes_list.append({
|
||||||
|
'id': batch_id,
|
||||||
|
'label': batch_info['name'], # 保持原始表述,如"国家级第1批"
|
||||||
|
'type': 'Batch',
|
||||||
|
'properties': json.dumps({
|
||||||
|
'original_string': batch_info['original_string']
|
||||||
|
}, ensure_ascii=False)
|
||||||
|
})
|
||||||
|
|
||||||
|
# 4. Inheritor节点(动态识别,保持原始表述)
|
||||||
|
print(" - 构建Inheritor节点...")
|
||||||
|
inheritors = extract_inheritors(df)
|
||||||
|
for inheritor_id, inheritor_info in inheritors.items():
|
||||||
|
nodes_list.append({
|
||||||
|
'id': inheritor_id,
|
||||||
|
'label': inheritor_info['name'], # 保持原始表述,如"吴明新(国)"
|
||||||
|
'type': 'Inheritor',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
# 5. Institution节点(动态识别,保持原始表述)
|
||||||
|
print(" - 构建Institution节点...")
|
||||||
|
institutions = extract_institutions(df)
|
||||||
|
for inst_id, inst_info in institutions.items():
|
||||||
|
nodes_list.append({
|
||||||
|
'id': inst_id,
|
||||||
|
'label': inst_info['name'], # 保持原始表述
|
||||||
|
'type': 'Institution',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
return nodes_list, batches, inheritors, institutions
|
||||||
|
|
||||||
|
|
||||||
|
def build_relations(df, batches, inheritors, institutions):
|
||||||
|
"""构建关系数据"""
|
||||||
|
relations_list = []
|
||||||
|
|
||||||
|
print("正在构建关系...")
|
||||||
|
|
||||||
|
# 建立名称到ID的快速查找映射
|
||||||
|
category_map = {}
|
||||||
|
for cat_name in df['类别'].dropna().unique():
|
||||||
|
cat_id = f"CAT-{abs(hash(cat_name)) % 100000:05d}"
|
||||||
|
category_map[cat_name] = cat_id
|
||||||
|
|
||||||
|
for idx, row in df.iterrows():
|
||||||
|
seq_num = row.get('总序号', idx + 1)
|
||||||
|
project_id = f"ICH-{int(seq_num)}"
|
||||||
|
|
||||||
|
# 1. 项目 → 类别
|
||||||
|
category_name = row.get('类别', '')
|
||||||
|
if category_name and category_name in category_map:
|
||||||
|
category_id = category_map[category_name]
|
||||||
|
relations_list.append({
|
||||||
|
'source': project_id,
|
||||||
|
'target': category_id,
|
||||||
|
'type': 'BELONGS_TO',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
# 2. 项目 → 批次(支持多个批次)
|
||||||
|
batch_str = row.get('项目批次', '')
|
||||||
|
if batch_str:
|
||||||
|
batch_ids = parse_batch_field(batch_str, batches)
|
||||||
|
for batch_id in batch_ids:
|
||||||
|
relations_list.append({
|
||||||
|
'source': project_id,
|
||||||
|
'target': batch_id,
|
||||||
|
'type': 'SELECTED_IN_BATCH',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
# 3. 项目 → 保护机构
|
||||||
|
institution_name = row.get('项目保护单位', '')
|
||||||
|
if institution_name and str(institution_name).strip() not in ['', '无']:
|
||||||
|
institution_id = get_institution_id(institution_name, institutions)
|
||||||
|
if institution_id:
|
||||||
|
relations_list.append({
|
||||||
|
'source': project_id,
|
||||||
|
'target': institution_id,
|
||||||
|
'type': 'PROTECTED_BY',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
# 4. 项目 → 传承人(支持多个传承人)
|
||||||
|
inheritor_str = row.get('代表性传承人', '')
|
||||||
|
if inheritor_str:
|
||||||
|
inheritor_names = parse_inheritor_field(inheritor_str)
|
||||||
|
for name in inheritor_names:
|
||||||
|
inheritor_id = get_inheritor_id(name, inheritors)
|
||||||
|
if inheritor_id:
|
||||||
|
relations_list.append({
|
||||||
|
'source': project_id,
|
||||||
|
'target': inheritor_id,
|
||||||
|
'type': 'HAS_INHERITOR',
|
||||||
|
'properties': '{}'
|
||||||
|
})
|
||||||
|
|
||||||
|
return relations_list
|
||||||
|
|
||||||
|
|
||||||
|
def generate_report(nodes_df, relations_df, output_dir):
|
||||||
|
"""生成统计报告"""
|
||||||
|
report_lines = []
|
||||||
|
report_lines.append("# 知识图谱CSV提取报告\n")
|
||||||
|
report_lines.append(f"生成时间: {pd.Timestamp.now().strftime('%Y-%m-%d %H:%M:%S')}\n")
|
||||||
|
report_lines.append("---\n\n")
|
||||||
|
|
||||||
|
# 节点统计
|
||||||
|
report_lines.append("## 节点统计\n\n")
|
||||||
|
report_lines.append(f"**节点总数**: {len(nodes_df)}\n\n")
|
||||||
|
|
||||||
|
node_type_counts = nodes_df['type'].value_counts().sort_index()
|
||||||
|
report_lines.append("| 节点类型 | 数量 | 占比 |\n")
|
||||||
|
report_lines.append("|---------|------|------|\n")
|
||||||
|
for node_type, count in node_type_counts.items():
|
||||||
|
percentage = (count / len(nodes_df) * 100)
|
||||||
|
report_lines.append(f"| {node_type} | {count} | {percentage:.1f}% |\n")
|
||||||
|
|
||||||
|
# 关系统计
|
||||||
|
report_lines.append("\n## 关系统计\n\n")
|
||||||
|
report_lines.append(f"**关系总数**: {len(relations_df)}\n\n")
|
||||||
|
|
||||||
|
rel_type_counts = relations_df['type'].value_counts().sort_index()
|
||||||
|
report_lines.append("| 关系类型 | 数量 | 占比 |\n")
|
||||||
|
report_lines.append("|---------|------|------|\n")
|
||||||
|
for rel_type, count in rel_type_counts.items():
|
||||||
|
percentage = (count / len(relations_df) * 100)
|
||||||
|
report_lines.append(f"| {rel_type} | {count} | {percentage:.1f}% |\n")
|
||||||
|
|
||||||
|
# 保存报告
|
||||||
|
report_path = output_dir / 'extraction_report.md'
|
||||||
|
with open(report_path, 'w', encoding='utf-8') as f:
|
||||||
|
f.writelines(report_lines)
|
||||||
|
|
||||||
|
print(f"\n报告已生成: {report_path}")
|
||||||
|
|
||||||
|
# 打印统计信息
|
||||||
|
print("\n" + "="*50)
|
||||||
|
print("数据提取完成")
|
||||||
|
print("="*50)
|
||||||
|
print(f"\n节点总数: {len(nodes_df)}")
|
||||||
|
for node_type, count in node_type_counts.items():
|
||||||
|
print(f" - {node_type}: {count}")
|
||||||
|
print(f"\n关系总数: {len(relations_df)}")
|
||||||
|
for rel_type, count in rel_type_counts.items():
|
||||||
|
print(f" - {rel_type}: {count}")
|
||||||
|
print("="*50)
|
||||||
|
|
||||||
|
|
||||||
|
def extract_data_from_excel():
|
||||||
|
"""从Excel提取数据并生成知识图谱CSV文件"""
|
||||||
|
|
||||||
|
print("开始从Excel提取数据...")
|
||||||
|
|
||||||
|
# 1. 读取Excel
|
||||||
|
excel_path = r'E:\Project\2026_KG_ICH\data\黑龙江国家级和省级非遗名单.xlsx'
|
||||||
|
print(f"读取文件: {excel_path}")
|
||||||
|
|
||||||
|
df = pd.read_excel(excel_path)
|
||||||
|
print(f"原始数据: {len(df)} 行 x {len(df.columns)} 列")
|
||||||
|
|
||||||
|
# 2. 提取所需列
|
||||||
|
columns = ['总序号', '类别', '项目名称', '项目批次', '项目保护单位', '代表性传承人']
|
||||||
|
print(f"\n提取列: {', '.join(columns)}")
|
||||||
|
|
||||||
|
# 检查列是否存在
|
||||||
|
available_cols = [col for col in columns if col in df.columns]
|
||||||
|
if len(available_cols) < len(columns):
|
||||||
|
missing = set(columns) - set(available_cols)
|
||||||
|
print(f"警告: 以下列不存在: {missing}")
|
||||||
|
|
||||||
|
df = df[available_cols]
|
||||||
|
print(f"提取后数据: {len(df)} 行 x {len(df.columns)} 列")
|
||||||
|
|
||||||
|
# 3. 数据清洗
|
||||||
|
print("\n数据清洗...")
|
||||||
|
df = clean_data(df)
|
||||||
|
print(f"清洗后数据: {len(df)} 行")
|
||||||
|
|
||||||
|
# 4. 构建节点
|
||||||
|
nodes_list, batches, inheritors, institutions = build_nodes(df)
|
||||||
|
nodes_df = pd.DataFrame(nodes_list)
|
||||||
|
print(f"节点数据: {len(nodes_df)} 行")
|
||||||
|
|
||||||
|
# 5. 构建关系
|
||||||
|
relations_list = build_relations(df, batches, inheritors, institutions)
|
||||||
|
relations_df = pd.DataFrame(relations_list)
|
||||||
|
print(f"关系数据: {len(relations_df)} 行")
|
||||||
|
|
||||||
|
# 6. 保存CSV
|
||||||
|
output_dir = Path(r'E:\Project\2026_KG_ICH\dofile\kg_project\output')
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
print(f"\n保存文件到: {output_dir}")
|
||||||
|
|
||||||
|
# 保存为UTF-8-BOM编码(Excel友好)
|
||||||
|
nodes_path = output_dir / 'nodes.csv'
|
||||||
|
relations_path = output_dir / 'rels.csv'
|
||||||
|
|
||||||
|
nodes_df.to_csv(nodes_path, index=False, encoding='utf-8-sig')
|
||||||
|
relations_df.to_csv(relations_path, index=False, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f" - nodes.csv: {nodes_path}")
|
||||||
|
print(f" - rels.csv: {relations_path}")
|
||||||
|
|
||||||
|
# 7. 生成统计报告
|
||||||
|
generate_report(nodes_df, relations_df, output_dir)
|
||||||
|
|
||||||
|
return nodes_df, relations_df
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
extract_data_from_excel()
|
||||||
@@ -0,0 +1,382 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
知识图谱可视化工具
|
||||||
|
使用networkx和matplotlib绘制知识图谱
|
||||||
|
"""
|
||||||
|
|
||||||
|
import pandas as pd
|
||||||
|
import networkx as nx
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
from matplotlib import font_manager
|
||||||
|
import matplotlib.patches as mpatches
|
||||||
|
from pathlib import Path
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
|
||||||
|
# 设置中文字体
|
||||||
|
def setup_chinese_font():
|
||||||
|
"""设置中文字体"""
|
||||||
|
# 尝试多种中文字体
|
||||||
|
chinese_fonts = [
|
||||||
|
'Microsoft YaHei',
|
||||||
|
'SimHei',
|
||||||
|
'SimSun',
|
||||||
|
'KaiTi',
|
||||||
|
'FangSong',
|
||||||
|
'STXihei',
|
||||||
|
'STSong',
|
||||||
|
'STKaiti',
|
||||||
|
'STFangsong'
|
||||||
|
]
|
||||||
|
|
||||||
|
for font in chinese_fonts:
|
||||||
|
try:
|
||||||
|
plt.rcParams['font.sans-serif'] = [font]
|
||||||
|
plt.rcParams['axes.unicode_minus'] = False
|
||||||
|
break
|
||||||
|
except:
|
||||||
|
continue
|
||||||
|
|
||||||
|
print(f"使用字体: {plt.rcParams['font.sans-serif'][0]}")
|
||||||
|
|
||||||
|
|
||||||
|
# 节点类型颜色映射
|
||||||
|
NODE_TYPE_COLORS = {
|
||||||
|
'ICH_Project': '#FF6B6B', # 红色 - 项目
|
||||||
|
'Category': '#4ECDC4', # 青色 - 类别
|
||||||
|
'Batch': '#95E1D3', # 绿色 - 批次
|
||||||
|
'Inheritor': '#FFD93D', # 黄色 - 传承人
|
||||||
|
'Institution': '#6C5CE7' # 紫色 - 机构
|
||||||
|
}
|
||||||
|
|
||||||
|
# 节点类型大小映射
|
||||||
|
NODE_TYPE_SIZES = {
|
||||||
|
'ICH_Project': 300,
|
||||||
|
'Category': 500,
|
||||||
|
'Batch': 350,
|
||||||
|
'Inheritor': 200,
|
||||||
|
'Institution': 250
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def load_graph_data(nodes_csv, rels_csv):
|
||||||
|
"""加载图谱数据"""
|
||||||
|
print("正在加载数据...")
|
||||||
|
|
||||||
|
# 读取节点和关系
|
||||||
|
nodes_df = pd.read_csv(nodes_csv, encoding='utf-8-sig')
|
||||||
|
rels_df = pd.read_csv(rels_csv, encoding='utf-8-sig')
|
||||||
|
|
||||||
|
print(f" - 节点: {len(nodes_df)}")
|
||||||
|
print(f" - 关系: {len(rels_df)}")
|
||||||
|
|
||||||
|
return nodes_df, rels_df
|
||||||
|
|
||||||
|
|
||||||
|
def build_networkx_graph(nodes_df, rels_df):
|
||||||
|
"""构建NetworkX图"""
|
||||||
|
print("正在构建图...")
|
||||||
|
|
||||||
|
G = nx.DiGraph() # 有向图
|
||||||
|
|
||||||
|
# 添加节点
|
||||||
|
for idx, row in nodes_df.iterrows():
|
||||||
|
node_id = row['id']
|
||||||
|
label = row['label']
|
||||||
|
node_type = row['type']
|
||||||
|
|
||||||
|
# 截断过长的标签
|
||||||
|
if len(label) > 10:
|
||||||
|
display_label = label[:10] + '...'
|
||||||
|
else:
|
||||||
|
display_label = label
|
||||||
|
|
||||||
|
G.add_node(
|
||||||
|
node_id,
|
||||||
|
label=display_label,
|
||||||
|
full_label=label,
|
||||||
|
node_type=node_type,
|
||||||
|
color=NODE_TYPE_COLORS.get(node_type, '#CCCCCC'),
|
||||||
|
size=NODE_TYPE_SIZES.get(node_type, 200)
|
||||||
|
)
|
||||||
|
|
||||||
|
# 添加边
|
||||||
|
for idx, row in rels_df.iterrows():
|
||||||
|
source = row['source']
|
||||||
|
target = row['target']
|
||||||
|
rel_type = row['type']
|
||||||
|
|
||||||
|
if source in G.nodes() and target in G.nodes():
|
||||||
|
G.add_edge(source, target, rel_type=rel_type)
|
||||||
|
|
||||||
|
print(f" - 节点数: {G.number_of_nodes()}")
|
||||||
|
print(f" - 边数: {G.number_of_edges()}")
|
||||||
|
|
||||||
|
return G
|
||||||
|
|
||||||
|
|
||||||
|
def filter_graph_by_type(G, include_types=None):
|
||||||
|
"""按节点类型过滤图"""
|
||||||
|
if include_types is None:
|
||||||
|
return G
|
||||||
|
|
||||||
|
nodes_to_keep = [n for n, d in G.nodes(data=True)
|
||||||
|
if d.get('node_type') in include_types]
|
||||||
|
|
||||||
|
return G.subgraph(nodes_to_keep).copy()
|
||||||
|
|
||||||
|
|
||||||
|
def draw_graph(G, output_path, title="知识图谱", layout='spring'):
|
||||||
|
"""绘制知识图谱"""
|
||||||
|
print(f"正在绘制图谱: {title}")
|
||||||
|
|
||||||
|
plt.figure(figsize=(20, 16))
|
||||||
|
|
||||||
|
# 选择布局算法
|
||||||
|
if layout == 'spring':
|
||||||
|
pos = nx.spring_layout(G, k=2, iterations=50, seed=42)
|
||||||
|
elif layout == 'circular':
|
||||||
|
pos = nx.circular_layout(G)
|
||||||
|
elif layout == 'kamada_kawai':
|
||||||
|
pos = nx.kamada_kawai_layout(G)
|
||||||
|
elif layout == 'random':
|
||||||
|
pos = nx.random_layout(G)
|
||||||
|
else:
|
||||||
|
pos = nx.spring_layout(G, k=2, iterations=50, seed=42)
|
||||||
|
|
||||||
|
# 按节点类型分组
|
||||||
|
node_types = {}
|
||||||
|
for node, data in G.nodes(data=True):
|
||||||
|
node_type = data.get('node_type', 'Unknown')
|
||||||
|
if node_type not in node_types:
|
||||||
|
node_types[node_type] = []
|
||||||
|
node_types[node_type].append(node)
|
||||||
|
|
||||||
|
# 绘制边
|
||||||
|
nx.draw_networkx_edges(
|
||||||
|
G, pos,
|
||||||
|
alpha=0.3,
|
||||||
|
width=0.5,
|
||||||
|
edge_color='gray',
|
||||||
|
arrows=True,
|
||||||
|
arrowsize=10,
|
||||||
|
arrowstyle='->,head_width=0.2,head_length=0.3'
|
||||||
|
)
|
||||||
|
|
||||||
|
# 按类型绘制节点
|
||||||
|
for node_type, nodes in node_types.items():
|
||||||
|
color = NODE_TYPE_COLORS.get(node_type, '#CCCCCC')
|
||||||
|
size = NODE_TYPE_SIZES.get(node_type, 200)
|
||||||
|
|
||||||
|
nx.draw_networkx_nodes(
|
||||||
|
G, pos,
|
||||||
|
nodelist=nodes,
|
||||||
|
node_color=color,
|
||||||
|
node_size=size,
|
||||||
|
alpha=0.8,
|
||||||
|
edgecolors='white',
|
||||||
|
linewidths=1
|
||||||
|
)
|
||||||
|
|
||||||
|
# 绘制标签(只对重要节点)
|
||||||
|
important_nodes = []
|
||||||
|
important_labels = {}
|
||||||
|
|
||||||
|
for node, data in G.nodes(data=True):
|
||||||
|
node_type = data.get('node_type')
|
||||||
|
# 只显示类别、批次和部分重要节点的标签
|
||||||
|
if node_type in ['Category', 'Batch'] or (
|
||||||
|
node_type == 'ICH_Project' and data.get('size', 0) > 400
|
||||||
|
):
|
||||||
|
important_nodes.append(node)
|
||||||
|
important_labels[node] = data.get('label', node)
|
||||||
|
|
||||||
|
if len(important_nodes) <= 100: # 节点不多时显示所有标签
|
||||||
|
nx.draw_networkx_labels(
|
||||||
|
G, pos,
|
||||||
|
labels=important_labels,
|
||||||
|
font_size=8,
|
||||||
|
font_weight='bold',
|
||||||
|
font_family='sans-serif'
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
# 节点太多时只显示类别标签
|
||||||
|
category_labels = {n: d['label'] for n, d in G.nodes(data=True)
|
||||||
|
if d.get('node_type') == 'Category'}
|
||||||
|
nx.draw_networkx_labels(
|
||||||
|
G, pos,
|
||||||
|
labels=category_labels,
|
||||||
|
font_size=10,
|
||||||
|
font_weight='bold'
|
||||||
|
)
|
||||||
|
|
||||||
|
# 图例
|
||||||
|
legend_patches = []
|
||||||
|
for node_type, color in NODE_TYPE_COLORS.items():
|
||||||
|
if node_type in node_types:
|
||||||
|
patch = mpatches.Patch(color=color, label=node_type)
|
||||||
|
legend_patches.append(patch)
|
||||||
|
|
||||||
|
plt.legend(
|
||||||
|
handles=legend_patches,
|
||||||
|
loc='upper right',
|
||||||
|
fontsize=12,
|
||||||
|
framealpha=0.9
|
||||||
|
)
|
||||||
|
|
||||||
|
plt.title(title, fontsize=16, fontweight='bold', pad=20)
|
||||||
|
plt.axis('off')
|
||||||
|
plt.tight_layout()
|
||||||
|
|
||||||
|
# 保存图片
|
||||||
|
plt.savefig(output_path, dpi=150, bbox_inches='tight')
|
||||||
|
print(f" - 保存到: {output_path}")
|
||||||
|
plt.close()
|
||||||
|
|
||||||
|
|
||||||
|
def draw_subgraphs(G, output_dir):
|
||||||
|
"""绘制子图(按节点类型分组)"""
|
||||||
|
print("\n正在绘制子图...")
|
||||||
|
|
||||||
|
output_dir = Path(output_dir)
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# 1. 只显示项目和类别
|
||||||
|
print(" 1. 项目-类别关系图...")
|
||||||
|
G1 = filter_graph_by_type(G, ['ICH_Project', 'Category'])
|
||||||
|
draw_graph(
|
||||||
|
G1,
|
||||||
|
output_dir / 'kg_project_category.png',
|
||||||
|
title='非遗项目与类别关系',
|
||||||
|
layout='spring'
|
||||||
|
)
|
||||||
|
|
||||||
|
# 2. 只显示项目和传承人
|
||||||
|
print(" 2. 项目-传承人关系图...")
|
||||||
|
G2 = filter_graph_by_type(G, ['ICH_Project', 'Inheritor'])
|
||||||
|
if G2.number_of_nodes() > 0:
|
||||||
|
draw_graph(
|
||||||
|
G2,
|
||||||
|
output_dir / 'kg_project_inheritor.png',
|
||||||
|
title='非遗项目与传承人关系',
|
||||||
|
layout='spring'
|
||||||
|
)
|
||||||
|
|
||||||
|
# 3. 只显示项目、类别和批次
|
||||||
|
print(" 3. 项目-类别-批次关系图...")
|
||||||
|
G3 = filter_graph_by_type(G, ['ICH_Project', 'Category', 'Batch'])
|
||||||
|
draw_graph(
|
||||||
|
G3,
|
||||||
|
output_dir / 'kg_project_category_batch.png',
|
||||||
|
title='非遗项目、类别与批次关系',
|
||||||
|
layout='kamada_kawai'
|
||||||
|
)
|
||||||
|
|
||||||
|
# 4. 完整图谱(抽样显示)
|
||||||
|
print(" 4. 完整知识图谱...")
|
||||||
|
if G.number_of_nodes() > 500:
|
||||||
|
# 节点太多时,只显示连接度高的节点
|
||||||
|
degrees = dict(G.degree())
|
||||||
|
high_degree_nodes = [n for n, d in degrees.items() if d >= 3]
|
||||||
|
G_sample = G.subgraph(high_degree_nodes).copy()
|
||||||
|
draw_graph(
|
||||||
|
G_sample,
|
||||||
|
output_dir / 'kg_full_sampled.png',
|
||||||
|
title=f'完整知识图谱(抽样,显示{G_sample.number_of_nodes()}个节点)',
|
||||||
|
layout='spring'
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
draw_graph(
|
||||||
|
G,
|
||||||
|
output_dir / 'kg_full.png',
|
||||||
|
title='完整知识图谱',
|
||||||
|
layout='spring'
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def print_statistics(G):
|
||||||
|
"""打印图统计信息"""
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("图谱统计信息")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
print(f"\n节点总数: {G.number_of_nodes()}")
|
||||||
|
print(f"边总数: {G.number_of_edges()}")
|
||||||
|
|
||||||
|
# 按类型统计节点
|
||||||
|
print("\n节点类型分布:")
|
||||||
|
node_types = {}
|
||||||
|
for node, data in G.nodes(data=True):
|
||||||
|
node_type = data.get('node_type', 'Unknown')
|
||||||
|
node_types[node_type] = node_types.get(node_type, 0) + 1
|
||||||
|
|
||||||
|
for node_type, count in sorted(node_types.items()):
|
||||||
|
percentage = (count / G.number_of_nodes() * 100)
|
||||||
|
print(f" - {node_type}: {count} ({percentage:.1f}%)")
|
||||||
|
|
||||||
|
# 按类型统计边
|
||||||
|
print("\n关系类型分布:")
|
||||||
|
rel_types = {}
|
||||||
|
for u, v, data in G.edges(data=True):
|
||||||
|
rel_type = data.get('rel_type', 'Unknown')
|
||||||
|
rel_types[rel_type] = rel_types.get(rel_type, 0) + 1
|
||||||
|
|
||||||
|
for rel_type, count in sorted(rel_types.items()):
|
||||||
|
percentage = (count / G.number_of_edges() * 100)
|
||||||
|
print(f" - {rel_type}: {count} ({percentage:.1f}%)")
|
||||||
|
|
||||||
|
# 连接度统计
|
||||||
|
degrees = [d for n, d in G.degree()]
|
||||||
|
print(f"\n连接度统计:")
|
||||||
|
print(f" - 平均连接度: {np.mean(degrees):.2f}")
|
||||||
|
print(f" - 最大连接度: {max(degrees)}")
|
||||||
|
print(f" - 最小连接度: {min(degrees)}")
|
||||||
|
|
||||||
|
# 找出连接度最高的节点
|
||||||
|
top_nodes = sorted(G.degree(), key=lambda x: x[1], reverse=True)[:10]
|
||||||
|
print(f"\n连接度最高的10个节点:")
|
||||||
|
for node, degree in top_nodes:
|
||||||
|
node_data = G.nodes[node]
|
||||||
|
label = node_data.get('full_label', node)
|
||||||
|
node_type = node_data.get('node_type', '')
|
||||||
|
print(f" - [{node_type}] {label}: {degree}个连接")
|
||||||
|
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
|
||||||
|
def visualize_kg(nodes_csv, rels_csv, output_dir):
|
||||||
|
"""可视化知识图谱"""
|
||||||
|
print("="*60)
|
||||||
|
print("知识图谱可视化工具")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
# 设置中文字体
|
||||||
|
setup_chinese_font()
|
||||||
|
|
||||||
|
# 加载数据
|
||||||
|
nodes_df, rels_df = load_graph_data(nodes_csv, rels_csv)
|
||||||
|
|
||||||
|
# 构建图
|
||||||
|
G = build_networkx_graph(nodes_df, rels_df)
|
||||||
|
|
||||||
|
# 打印统计信息
|
||||||
|
print_statistics(G)
|
||||||
|
|
||||||
|
# 绘制子图
|
||||||
|
draw_subgraphs(G, output_dir)
|
||||||
|
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("可视化完成!")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
# 输入文件
|
||||||
|
nodes_csv = r'E:\Project\2026_KG_ICH\dofile\kg_project\output\nodes.csv'
|
||||||
|
rels_csv = r'E:\Project\2026_KG_ICH\dofile\kg_project\output\rels.csv'
|
||||||
|
|
||||||
|
# 输出目录
|
||||||
|
output_dir = r'E:\Project\2026_KG_ICH\dofile\kg_project\output\visualizations'
|
||||||
|
|
||||||
|
# 执行可视化
|
||||||
|
visualize_kg(nodes_csv, rels_csv, output_dir)
|
||||||
@@ -0,0 +1,344 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
实体规范化器
|
||||||
|
负责实体的去重、ID生成和规范化
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import hashlib
|
||||||
|
import logging
|
||||||
|
from typing import Dict, List, Any, Optional
|
||||||
|
from pathlib import Path
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
|
||||||
|
class EntityNormalizer:
|
||||||
|
"""实体规范化器"""
|
||||||
|
|
||||||
|
def __init__(self, ontology_file: str):
|
||||||
|
"""
|
||||||
|
初始化规范化器
|
||||||
|
|
||||||
|
Args:
|
||||||
|
ontology_file: 本体配置文件路径
|
||||||
|
"""
|
||||||
|
# 加载本体配置
|
||||||
|
with open(ontology_file, 'r', encoding='utf-8') as f:
|
||||||
|
self.ontology = yaml.safe_load(f)
|
||||||
|
|
||||||
|
self.logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# 实体注册表:{entity_id: entity_data}
|
||||||
|
self.entity_registry = {}
|
||||||
|
|
||||||
|
# 文本到ID的映射:{normalized_text: entity_id}
|
||||||
|
self.text_to_id_map = {}
|
||||||
|
|
||||||
|
# 类型计数器:{entity_type: count}
|
||||||
|
self.type_counters = {}
|
||||||
|
|
||||||
|
self.logger.info("EntityNormalizer初始化完成")
|
||||||
|
|
||||||
|
def normalize_text(self, text: str) -> str:
|
||||||
|
"""
|
||||||
|
标准化文本
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text: 原始文本
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
标准化后的文本
|
||||||
|
"""
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
|
||||||
|
# 去除首尾空格
|
||||||
|
text = text.strip()
|
||||||
|
|
||||||
|
# 统一全角/半角字符
|
||||||
|
text = text.replace(' ', ' ').replace(',', ',').replace('、', ',')
|
||||||
|
|
||||||
|
# 转换为小写进行比较(保持原文用于显示)
|
||||||
|
return text
|
||||||
|
|
||||||
|
def calculate_similarity(self, text1: str, text2: str) -> float:
|
||||||
|
"""
|
||||||
|
计算两个文本的相似度(基于编辑距离)
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text1: 文本1
|
||||||
|
text2: 文本2
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
相似度(0-1之间)
|
||||||
|
"""
|
||||||
|
import Levenshtein
|
||||||
|
|
||||||
|
norm1 = self.normalize_text(text1)
|
||||||
|
norm2 = self.normalize_text(text2)
|
||||||
|
|
||||||
|
if not norm1 or not norm2:
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
max_len = max(len(norm1), len(norm2))
|
||||||
|
if max_len == 0:
|
||||||
|
return 1.0
|
||||||
|
|
||||||
|
distance = Levenshtein.distance(norm1, norm2)
|
||||||
|
similarity = 1.0 - (distance / max_len)
|
||||||
|
|
||||||
|
return similarity
|
||||||
|
|
||||||
|
def generate_entity_id(self, entity_type: str, entity_text: str) -> str:
|
||||||
|
"""
|
||||||
|
生成实体ID
|
||||||
|
|
||||||
|
Args:
|
||||||
|
entity_type: 实体类型
|
||||||
|
entity_text: 实体文本
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
实体ID
|
||||||
|
"""
|
||||||
|
# 获取类型配置
|
||||||
|
type_config = self.ontology['entity_types'].get(entity_type, {})
|
||||||
|
prefix = type_config.get('prefix', 'UNK')
|
||||||
|
|
||||||
|
# 生成哈希值
|
||||||
|
text_hash = abs(hash(entity_text)) % 100000
|
||||||
|
|
||||||
|
# 检查是否需要计数器(某些类型可能需要)
|
||||||
|
if entity_type not in self.type_counters:
|
||||||
|
self.type_counters[entity_type] = 0
|
||||||
|
|
||||||
|
# 根据类型生成ID
|
||||||
|
if entity_type in ['Ethnic_Group', 'Geographic_Location', 'Geographic_Environment', 'Time_Period']:
|
||||||
|
# 这些类型使用唯一哈希,不需要计数器
|
||||||
|
entity_id = f"{prefix}-{text_hash:05d}"
|
||||||
|
else:
|
||||||
|
# 其他类型使用计数器
|
||||||
|
self.type_counters[entity_type] += 1
|
||||||
|
counter = self.type_counters[entity_type]
|
||||||
|
entity_id = f"{prefix}-{text_hash:05d}-{counter}"
|
||||||
|
|
||||||
|
return entity_id
|
||||||
|
|
||||||
|
def find_similar_entity(
|
||||||
|
self,
|
||||||
|
entity_text: str,
|
||||||
|
entity_type: str,
|
||||||
|
similarity_threshold: float = 0.85
|
||||||
|
) -> Optional[str]:
|
||||||
|
"""
|
||||||
|
查找相似实体
|
||||||
|
|
||||||
|
Args:
|
||||||
|
entity_text: 实体文本
|
||||||
|
entity_type: 实体类型
|
||||||
|
similarity_threshold: 相似度阈值
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
相似实体的ID,如果不存在返回None
|
||||||
|
"""
|
||||||
|
normalized_text = self.normalize_text(entity_text)
|
||||||
|
|
||||||
|
# 只在相同类型的实体中查找
|
||||||
|
for entity_id, entity_data in self.entity_registry.items():
|
||||||
|
if entity_data['type'] != entity_type:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 检查文本相似度
|
||||||
|
similarity = self.calculate_similarity(normalized_text, entity_data['canonical_name'])
|
||||||
|
|
||||||
|
if similarity >= similarity_threshold:
|
||||||
|
self.logger.info(f"发现相似实体: '{entity_text}' ~ '{entity_data['canonical_name']}' (相似度: {similarity:.2f})")
|
||||||
|
return entity_id
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
def is_generic_concept(self, entity_text: str, entity_type: str) -> bool:
|
||||||
|
"""
|
||||||
|
检查是否为泛指概念(应该被过滤)
|
||||||
|
|
||||||
|
Args:
|
||||||
|
entity_text: 实体文本
|
||||||
|
entity_type: 实体类型
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
True表示是泛指概念,应该过滤
|
||||||
|
"""
|
||||||
|
generic_patterns = [
|
||||||
|
'东北少数民族',
|
||||||
|
'本地土著民族',
|
||||||
|
'当地民族',
|
||||||
|
'少数民族',
|
||||||
|
'土著',
|
||||||
|
'东北民族',
|
||||||
|
'本地民族',
|
||||||
|
'地区民族',
|
||||||
|
]
|
||||||
|
|
||||||
|
normalized_text = self.normalize_text(entity_text)
|
||||||
|
|
||||||
|
# 检查是否匹配泛指模式
|
||||||
|
for pattern in generic_patterns:
|
||||||
|
if pattern in normalized_text:
|
||||||
|
self.logger.info(f"过滤泛指概念: '{entity_text}' (类型: {entity_type})")
|
||||||
|
return True
|
||||||
|
|
||||||
|
# 对于Ethnic_Group类型,必须是具体的民族名称
|
||||||
|
if entity_type == 'Ethnic_Group':
|
||||||
|
specific_ethnic_groups = [
|
||||||
|
'满族', '赫哲族', '鄂伦春族', '鄂温克族',
|
||||||
|
'达斡尔族', '朝鲜族', '蒙古族', '回族',
|
||||||
|
'汉族', '锡伯族', '柯尔克孜族'
|
||||||
|
]
|
||||||
|
# 如果不在具体民族列表中,可能是泛指概念
|
||||||
|
is_specific = any(group in normalized_text for group in specific_ethnic_groups)
|
||||||
|
if not is_specific:
|
||||||
|
self.logger.info(f"过滤非具体民族: '{entity_text}'")
|
||||||
|
return True
|
||||||
|
|
||||||
|
return False
|
||||||
|
|
||||||
|
def normalize_entity(
|
||||||
|
self,
|
||||||
|
raw_entity: Dict[str, Any],
|
||||||
|
similarity_threshold: float = 0.85,
|
||||||
|
filter_generic: bool = True
|
||||||
|
) -> str:
|
||||||
|
"""
|
||||||
|
规范化单个实体
|
||||||
|
|
||||||
|
Args:
|
||||||
|
raw_entity: 原始实体数据(从LLM返回)
|
||||||
|
similarity_threshold: 相似度阈值
|
||||||
|
filter_generic: 是否过滤泛指概念
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
实体ID
|
||||||
|
"""
|
||||||
|
entity_text = raw_entity.get('text', '')
|
||||||
|
entity_type = raw_entity.get('type', '')
|
||||||
|
attributes = raw_entity.get('attributes', {})
|
||||||
|
|
||||||
|
if not entity_text or not entity_type:
|
||||||
|
self.logger.error(f"实体缺少text或type: {raw_entity}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 验证实体类型是否在本体中定义
|
||||||
|
if entity_type not in self.ontology['entity_types']:
|
||||||
|
self.logger.warning(f"未知实体类型 '{entity_type}',跳过: '{entity_text}'")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 过滤泛指概念
|
||||||
|
if filter_generic and self.is_generic_concept(entity_text, entity_type):
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 标准化文本
|
||||||
|
normalized_text = self.normalize_text(entity_text)
|
||||||
|
|
||||||
|
# 检查是否已存在完全相同的实体
|
||||||
|
if normalized_text in self.text_to_id_map:
|
||||||
|
existing_id = self.text_to_id_map[normalized_text]
|
||||||
|
self.logger.info(f"实体已存在: '{entity_text}' -> {existing_id}")
|
||||||
|
return existing_id
|
||||||
|
|
||||||
|
# 查找相似实体
|
||||||
|
similar_id = self.find_similar_entity(entity_text, entity_type, similarity_threshold)
|
||||||
|
if similar_id:
|
||||||
|
# 合并到相似实体
|
||||||
|
self.logger.info(f"合并实体: '{entity_text}' -> {similar_id}")
|
||||||
|
self.text_to_id_map[normalized_text] = similar_id
|
||||||
|
return similar_id
|
||||||
|
|
||||||
|
# 生成新实体ID
|
||||||
|
entity_id = self.generate_entity_id(entity_type, entity_text)
|
||||||
|
|
||||||
|
# 获取标准名称
|
||||||
|
standard_name = attributes.get('name') or entity_text
|
||||||
|
|
||||||
|
# 保存到注册表
|
||||||
|
self.entity_registry[entity_id] = {
|
||||||
|
'canonical_name': normalized_text,
|
||||||
|
'display_name': standard_name,
|
||||||
|
'type': entity_type,
|
||||||
|
'attributes': attributes,
|
||||||
|
'source_projects': []
|
||||||
|
}
|
||||||
|
|
||||||
|
# 保存文本映射
|
||||||
|
self.text_to_id_map[normalized_text] = entity_id
|
||||||
|
|
||||||
|
self.logger.info(f"新建实体: {entity_id} - '{standard_name}' ({entity_type})")
|
||||||
|
|
||||||
|
return entity_id
|
||||||
|
|
||||||
|
def normalize_batch(
|
||||||
|
self,
|
||||||
|
extraction_results: List[Dict[str, Any]],
|
||||||
|
similarity_threshold: float = 0.85
|
||||||
|
) -> Dict[str, str]:
|
||||||
|
"""
|
||||||
|
批量规范化实体
|
||||||
|
|
||||||
|
Args:
|
||||||
|
extraction_results: LLM抽取结果列表
|
||||||
|
similarity_threshold: 相似度阈值
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
文本到ID的映射字典(包含所有已处理的实体)
|
||||||
|
"""
|
||||||
|
self.logger.info(f"开始批量规范化,共{len(extraction_results)}个抽取结果")
|
||||||
|
|
||||||
|
for idx, result in enumerate(extraction_results):
|
||||||
|
if not result or 'entities' not in result:
|
||||||
|
continue
|
||||||
|
|
||||||
|
for entity in result['entities']:
|
||||||
|
self.normalize_entity(entity, similarity_threshold)
|
||||||
|
|
||||||
|
self.logger.info(f"批量规范化完成,生成{len(self.entity_registry)}个唯一实体")
|
||||||
|
|
||||||
|
# 返回完整的文本到ID映射(包含所有已处理的实体,不仅仅是本批次)
|
||||||
|
return self.text_to_id_map.copy()
|
||||||
|
|
||||||
|
def get_entity_nodes(self) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
获取所有实体节点(用于生成CSV)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
节点列表
|
||||||
|
"""
|
||||||
|
nodes = []
|
||||||
|
|
||||||
|
for entity_id, entity_data in self.entity_registry.items():
|
||||||
|
node = {
|
||||||
|
'id': entity_id,
|
||||||
|
'label': entity_data['display_name'],
|
||||||
|
'type': entity_data['type'],
|
||||||
|
'properties': json.dumps(entity_data['attributes'], ensure_ascii=False)
|
||||||
|
}
|
||||||
|
nodes.append(node)
|
||||||
|
|
||||||
|
return nodes
|
||||||
|
|
||||||
|
def get_statistics(self) -> Dict[str, Any]:
|
||||||
|
"""
|
||||||
|
获取统计信息
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
统计信息字典
|
||||||
|
"""
|
||||||
|
stats = {
|
||||||
|
'total_entities': len(self.entity_registry),
|
||||||
|
'entities_by_type': {},
|
||||||
|
'type_counters': self.type_counters.copy()
|
||||||
|
}
|
||||||
|
|
||||||
|
# 按类型统计
|
||||||
|
for entity_data in self.entity_registry.values():
|
||||||
|
entity_type = entity_data['type']
|
||||||
|
stats['entities_by_type'][entity_type] = stats['entities_by_type'].get(entity_type, 0) + 1
|
||||||
|
|
||||||
|
return stats
|
||||||
@@ -0,0 +1,280 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
关系构建器
|
||||||
|
负责构建非遗项目与深层实体之间的关系
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
from typing import Dict, List, Any, Optional
|
||||||
|
from pathlib import Path
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
|
||||||
|
class RelationshipBuilder:
|
||||||
|
"""关系构建器"""
|
||||||
|
|
||||||
|
def __init__(self, ontology_file: str):
|
||||||
|
"""
|
||||||
|
初始化关系构建器
|
||||||
|
|
||||||
|
Args:
|
||||||
|
ontology_file: 本体配置文件路径
|
||||||
|
"""
|
||||||
|
# 加载本体配置
|
||||||
|
with open(ontology_file, 'r', encoding='utf-8') as f:
|
||||||
|
self.ontology = yaml.safe_load(f)
|
||||||
|
|
||||||
|
self.logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# 关系注册表:用于去重
|
||||||
|
self.relationship_registry = set()
|
||||||
|
|
||||||
|
self.logger.info("RelationshipBuilder初始化完成")
|
||||||
|
|
||||||
|
def build_relationship(
|
||||||
|
self,
|
||||||
|
source_id: str,
|
||||||
|
target_id: str,
|
||||||
|
rel_type: str,
|
||||||
|
properties: Dict[str, Any] = None
|
||||||
|
) -> Optional[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
构建单个关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
source_id: 源实体ID
|
||||||
|
target_id: 目标实体ID
|
||||||
|
rel_type: 关系类型
|
||||||
|
properties: 关系属性
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
关系字典,如果验证失败返回None
|
||||||
|
"""
|
||||||
|
# 验证关系类型
|
||||||
|
if rel_type not in self.ontology['relationship_types']:
|
||||||
|
self.logger.warning(f"未知关系类型: {rel_type}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 验证ID不为空
|
||||||
|
if not source_id or not target_id:
|
||||||
|
self.logger.error(f"关系ID不能为空: source={source_id}, target={target_id}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 生成关系唯一键(用于去重)
|
||||||
|
rel_key = f"{source_id}-{target_id}-{rel_type}"
|
||||||
|
|
||||||
|
# 检查是否重复
|
||||||
|
if rel_key in self.relationship_registry:
|
||||||
|
self.logger.debug(f"关系已存在,跳过: {rel_key}")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 添加到注册表
|
||||||
|
self.relationship_registry.add(rel_key)
|
||||||
|
|
||||||
|
# 构建关系
|
||||||
|
relationship = {
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type,
|
||||||
|
'properties': json.dumps(properties or {}, ensure_ascii=False)
|
||||||
|
}
|
||||||
|
|
||||||
|
return relationship
|
||||||
|
|
||||||
|
def build_relationships_from_extraction(
|
||||||
|
self,
|
||||||
|
project_id: str,
|
||||||
|
extraction_result: Dict[str, Any],
|
||||||
|
entity_id_map: Dict[str, str]
|
||||||
|
) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
从抽取结果构建关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
project_id: 项目ID
|
||||||
|
extraction_result: LLM抽取结果
|
||||||
|
entity_id_map: 实体文本到ID的映射
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
关系列表
|
||||||
|
"""
|
||||||
|
relationships = []
|
||||||
|
|
||||||
|
if not extraction_result or 'relationships' not in extraction_result:
|
||||||
|
return relationships
|
||||||
|
|
||||||
|
for rel in extraction_result['relationships']:
|
||||||
|
# 获取源实体(应该是项目ID)
|
||||||
|
source_entity = rel.get('source_entity', '')
|
||||||
|
|
||||||
|
# 验证源实体是否匹配当前项目
|
||||||
|
if source_entity != project_id:
|
||||||
|
self.logger.warning(f"关系源实体不匹配: 期望{project_id}, 实际{source_entity}")
|
||||||
|
# 如果不匹配,尝试修正
|
||||||
|
source_entity = project_id
|
||||||
|
|
||||||
|
# 获取目标实体文本
|
||||||
|
target_text = rel.get('target_entity', '')
|
||||||
|
|
||||||
|
# 查找目标实体ID
|
||||||
|
target_id = entity_id_map.get(target_text)
|
||||||
|
|
||||||
|
if not target_id:
|
||||||
|
self.logger.warning(f"未找到目标实体ID: {target_text}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 获取关系类型
|
||||||
|
rel_type = rel.get('type', '')
|
||||||
|
rel_properties = rel.get('properties', {})
|
||||||
|
|
||||||
|
# 构建关系
|
||||||
|
relationship = self.build_relationship(
|
||||||
|
source_entity,
|
||||||
|
target_id,
|
||||||
|
rel_type,
|
||||||
|
rel_properties
|
||||||
|
)
|
||||||
|
|
||||||
|
if relationship:
|
||||||
|
relationships.append(relationship)
|
||||||
|
|
||||||
|
return relationships
|
||||||
|
|
||||||
|
def build_batch_relationships(
|
||||||
|
self,
|
||||||
|
projects_data: List[Dict[str, Any]],
|
||||||
|
extraction_results: List[Dict[str, Any]],
|
||||||
|
entity_id_map: Dict[str, str]
|
||||||
|
) -> List[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
批量构建关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
projects_data: 项目数据列表
|
||||||
|
extraction_results: 抽取结果列表
|
||||||
|
entity_id_map: 实体文本到ID的映射
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
关系列表
|
||||||
|
"""
|
||||||
|
self.logger.info(f"开始批量构建关系,共{len(projects_data)}个项目")
|
||||||
|
|
||||||
|
all_relationships = []
|
||||||
|
|
||||||
|
for idx, (project_data, extraction_result) in enumerate(zip(projects_data, extraction_results)):
|
||||||
|
if not extraction_result:
|
||||||
|
continue
|
||||||
|
|
||||||
|
project_id = project_data.get('project_id', '')
|
||||||
|
|
||||||
|
if not project_id:
|
||||||
|
self.logger.warning(f"项目{idx}缺少project_id")
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 构建该项目的关系
|
||||||
|
relationships = self.build_relationships_from_extraction(
|
||||||
|
project_id,
|
||||||
|
extraction_result,
|
||||||
|
entity_id_map
|
||||||
|
)
|
||||||
|
|
||||||
|
all_relationships.extend(relationships)
|
||||||
|
|
||||||
|
self.logger.info(f"项目 {project_id} 构建了{len(relationships)}个关系")
|
||||||
|
|
||||||
|
self.logger.info(f"批量构建完成,共{len(all_relationships)}个关系")
|
||||||
|
|
||||||
|
return all_relationships
|
||||||
|
|
||||||
|
def validate_relationships(
|
||||||
|
self,
|
||||||
|
relationships: List[Dict[str, Any]],
|
||||||
|
node_ids: set
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""
|
||||||
|
验证关系完整性
|
||||||
|
|
||||||
|
Args:
|
||||||
|
relationships: 关系列表
|
||||||
|
node_ids: 节点ID集合
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
验证报告
|
||||||
|
"""
|
||||||
|
report = {
|
||||||
|
'total_relationships': len(relationships),
|
||||||
|
'valid_relationships': 0,
|
||||||
|
'invalid_relationships': 0,
|
||||||
|
'broken_links': [],
|
||||||
|
'invalid_types': [],
|
||||||
|
'relationships_by_type': {}
|
||||||
|
}
|
||||||
|
|
||||||
|
for rel in relationships:
|
||||||
|
source_id = rel.get('source', '')
|
||||||
|
target_id = rel.get('target', '')
|
||||||
|
rel_type = rel.get('type', '')
|
||||||
|
|
||||||
|
# 统计关系类型
|
||||||
|
report['relationships_by_type'][rel_type] = \
|
||||||
|
report['relationships_by_type'].get(rel_type, 0) + 1
|
||||||
|
|
||||||
|
is_valid = True
|
||||||
|
|
||||||
|
# 验证关系类型
|
||||||
|
if rel_type not in self.ontology['relationship_types']:
|
||||||
|
report['invalid_types'].append({
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type
|
||||||
|
})
|
||||||
|
is_valid = False
|
||||||
|
|
||||||
|
# 验证链接完整性
|
||||||
|
if source_id not in node_ids:
|
||||||
|
report['broken_links'].append({
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type,
|
||||||
|
'issue': 'source_not_found'
|
||||||
|
})
|
||||||
|
is_valid = False
|
||||||
|
|
||||||
|
if target_id not in node_ids:
|
||||||
|
report['broken_links'].append({
|
||||||
|
'source': source_id,
|
||||||
|
'target': target_id,
|
||||||
|
'type': rel_type,
|
||||||
|
'issue': 'target_not_found'
|
||||||
|
})
|
||||||
|
is_valid = False
|
||||||
|
|
||||||
|
if is_valid:
|
||||||
|
report['valid_relationships'] += 1
|
||||||
|
else:
|
||||||
|
report['invalid_relationships'] += 1
|
||||||
|
|
||||||
|
return report
|
||||||
|
|
||||||
|
def get_statistics(self, relationships: List[Dict[str, Any]]) -> Dict[str, Any]:
|
||||||
|
"""
|
||||||
|
获取关系统计信息
|
||||||
|
|
||||||
|
Args:
|
||||||
|
relationships: 关系列表
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
统计信息
|
||||||
|
"""
|
||||||
|
stats = {
|
||||||
|
'total_relationships': len(relationships),
|
||||||
|
'relationships_by_type': {}
|
||||||
|
}
|
||||||
|
|
||||||
|
for rel in relationships:
|
||||||
|
rel_type = rel.get('type', 'Unknown')
|
||||||
|
stats['relationships_by_type'][rel_type] = \
|
||||||
|
stats['relationships_by_type'].get(rel_type, 0) + 1
|
||||||
|
|
||||||
|
return stats
|
||||||
@@ -0,0 +1,406 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
深度文化实体抽取主流程
|
||||||
|
从Excel备注字段抽取深层文化实体并生成知识图谱CSV文件
|
||||||
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import pandas as pd
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
from pathlib import Path
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Dict, List, Any
|
||||||
|
|
||||||
|
# 添加模块路径
|
||||||
|
import sys
|
||||||
|
sys.path.append(str(Path(__file__).parent))
|
||||||
|
|
||||||
|
from knowledge_extraction.deep_entity_extractor import DeepEntityExtractor
|
||||||
|
from data_processing.entity_normalizer import EntityNormalizer
|
||||||
|
from data_processing.relationship_builder import RelationshipBuilder
|
||||||
|
|
||||||
|
|
||||||
|
class DeepExtractionPipeline:
|
||||||
|
"""深度文化实体抽取流程"""
|
||||||
|
|
||||||
|
def __init__(self, config_file: str):
|
||||||
|
"""
|
||||||
|
初始化流程
|
||||||
|
|
||||||
|
Args:
|
||||||
|
config_file: 配置文件路径
|
||||||
|
"""
|
||||||
|
self.config_file = config_file
|
||||||
|
|
||||||
|
# 设置日志
|
||||||
|
logging.basicConfig(
|
||||||
|
level=logging.INFO,
|
||||||
|
format='%(asctime)s - %(name)s - %(levelname)s - %(message)s'
|
||||||
|
)
|
||||||
|
self.logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# 初始化组件
|
||||||
|
self.extractor = DeepEntityExtractor(config_file)
|
||||||
|
self.normalizer = EntityNormalizer(
|
||||||
|
str(Path(config_file).parent / 'entity_ontology.yaml')
|
||||||
|
)
|
||||||
|
self.builder = RelationshipBuilder(
|
||||||
|
str(Path(config_file).parent / 'entity_ontology.yaml')
|
||||||
|
)
|
||||||
|
|
||||||
|
self.logger.info("DeepExtractionPipeline初始化完成")
|
||||||
|
|
||||||
|
def load_excel_data(self, excel_file: str, max_rows: int = None) -> List[Dict[str, str]]:
|
||||||
|
"""
|
||||||
|
加载Excel数据
|
||||||
|
|
||||||
|
Args:
|
||||||
|
excel_file: Excel文件路径
|
||||||
|
max_rows: 最大读取行数(用于测试)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
项目数据列表
|
||||||
|
"""
|
||||||
|
self.logger.info(f"读取Excel文件: {excel_file}")
|
||||||
|
|
||||||
|
df = pd.read_excel(excel_file)
|
||||||
|
|
||||||
|
# 限制行数
|
||||||
|
if max_rows:
|
||||||
|
df = df.head(max_rows)
|
||||||
|
self.logger.info(f"限制读取行数: {max_rows}")
|
||||||
|
|
||||||
|
# 检查必需列
|
||||||
|
required_columns = ['总序号', '项目名称', '备注']
|
||||||
|
missing_columns = [col for col in required_columns if col not in df.columns]
|
||||||
|
|
||||||
|
if missing_columns:
|
||||||
|
self.logger.error(f"Excel缺少必需列: {missing_columns}")
|
||||||
|
raise ValueError(f"缺少必需列: {missing_columns}")
|
||||||
|
|
||||||
|
# 构建项目数据列表
|
||||||
|
projects_data = []
|
||||||
|
for idx, row in df.iterrows():
|
||||||
|
seq_num = row.get('总序号', idx + 1)
|
||||||
|
project_name = row.get('项目名称', '')
|
||||||
|
remark_text = row.get('备注', '')
|
||||||
|
|
||||||
|
# 过滤空备注
|
||||||
|
if not remark_text or len(str(remark_text).strip()) < 10:
|
||||||
|
continue
|
||||||
|
|
||||||
|
projects_data.append({
|
||||||
|
'project_id': f"ICH-{int(seq_num)}",
|
||||||
|
'project_name': str(project_name),
|
||||||
|
'remark_text': str(remark_text),
|
||||||
|
'row_number': idx + 2 # Excel行号(1-based + header)
|
||||||
|
})
|
||||||
|
|
||||||
|
self.logger.info(f"加载了{len(projects_data)}个项目的数据")
|
||||||
|
|
||||||
|
return projects_data
|
||||||
|
|
||||||
|
async def run_extraction(
|
||||||
|
self,
|
||||||
|
excel_file: str,
|
||||||
|
output_prefix: str = "",
|
||||||
|
max_rows: int = None
|
||||||
|
) -> Dict[str, Any]:
|
||||||
|
"""
|
||||||
|
执行完整的抽取流程(支持增量保存)
|
||||||
|
|
||||||
|
Args:
|
||||||
|
excel_file: Excel文件路径
|
||||||
|
output_prefix: 输出文件前缀(用于测试)
|
||||||
|
max_rows: 最大读取行数(用于测试)
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
处理结果统计
|
||||||
|
"""
|
||||||
|
self.logger.info("="*60)
|
||||||
|
self.logger.info("开始深度文化实体抽取流程(增量保存模式)")
|
||||||
|
self.logger.info("="*60)
|
||||||
|
|
||||||
|
start_time = datetime.now()
|
||||||
|
|
||||||
|
# 1. 加载数据
|
||||||
|
projects_data = self.load_excel_data(excel_file, max_rows)
|
||||||
|
|
||||||
|
if not projects_data:
|
||||||
|
self.logger.error("没有可处理的数据")
|
||||||
|
return {}
|
||||||
|
|
||||||
|
# 准备输出目录
|
||||||
|
output_dir = Path('output')
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# 确定输出文件名
|
||||||
|
if output_prefix:
|
||||||
|
nodes_file = output_dir / f"{output_prefix}_nodes.csv"
|
||||||
|
rels_file = output_dir / f"{output_prefix}_rels.csv"
|
||||||
|
report_file = output_dir / f"{output_prefix}_report.md"
|
||||||
|
else:
|
||||||
|
nodes_file = output_dir / Path(self.extractor.config['output']['nodes_file']).name
|
||||||
|
rels_file = output_dir / Path(self.extractor.config['output']['relationships_file']).name
|
||||||
|
report_file = output_dir / Path(self.extractor.config['output']['report_file']).name
|
||||||
|
|
||||||
|
# 2. 定义增量保存回调函数
|
||||||
|
async def save_progress(batch_num, total_batches, extraction_results, processed_projects):
|
||||||
|
"""每批次完成后保存进度(仅保存节点,不构建关系)"""
|
||||||
|
try:
|
||||||
|
self.logger.info(f"\n[增量保存] 批次 {batch_num}/{total_batches} 开始保存...")
|
||||||
|
|
||||||
|
# 实体规范化(累积到实体注册表)
|
||||||
|
self.normalizer.normalize_batch(extraction_results)
|
||||||
|
|
||||||
|
# 生成节点(累积所有已处理的实体)
|
||||||
|
entity_nodes = self.normalizer.get_entity_nodes()
|
||||||
|
|
||||||
|
# 添加项目节点(仅当前批次)
|
||||||
|
project_nodes = [
|
||||||
|
{
|
||||||
|
'id': item['project_id'],
|
||||||
|
'label': item['project_name'],
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': '{}'
|
||||||
|
}
|
||||||
|
for item in processed_projects
|
||||||
|
]
|
||||||
|
|
||||||
|
all_nodes = entity_nodes + project_nodes
|
||||||
|
|
||||||
|
# 保存节点
|
||||||
|
nodes_df = pd.DataFrame(all_nodes)
|
||||||
|
nodes_df.to_csv(nodes_file, index=False, encoding='utf-8-sig')
|
||||||
|
self.logger.info(f"[增量保存] 节点已保存: {nodes_file} ({len(all_nodes)}行)")
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"[增量保存] 批次 {batch_num} 保存失败: {str(e)}", exc_info=True)
|
||||||
|
|
||||||
|
# 3. LLM批量抽取(带增量保存回调)
|
||||||
|
self.logger.info("\n步骤1: LLM批量抽取(增量保存模式)")
|
||||||
|
extraction_results = await self.extractor.batch_extract(
|
||||||
|
projects_data,
|
||||||
|
progress_callback=save_progress
|
||||||
|
)
|
||||||
|
|
||||||
|
# 4. 最终数据验证和报告
|
||||||
|
self.logger.info("\n步骤2: 最终数据验证")
|
||||||
|
|
||||||
|
# 重新规范化所有实体(确保一致性)
|
||||||
|
entity_id_map = self.normalizer.normalize_batch(extraction_results)
|
||||||
|
|
||||||
|
# 重新构建所有关系
|
||||||
|
relationships = self.builder.build_batch_relationships(
|
||||||
|
projects_data,
|
||||||
|
extraction_results,
|
||||||
|
entity_id_map
|
||||||
|
)
|
||||||
|
|
||||||
|
# 生成最终节点
|
||||||
|
entity_nodes = self.normalizer.get_entity_nodes()
|
||||||
|
project_nodes = [
|
||||||
|
{
|
||||||
|
'id': item['project_id'],
|
||||||
|
'label': item['project_name'],
|
||||||
|
'type': 'ICH_Project',
|
||||||
|
'properties': '{}'
|
||||||
|
}
|
||||||
|
for item in projects_data
|
||||||
|
]
|
||||||
|
all_nodes = entity_nodes + project_nodes
|
||||||
|
|
||||||
|
# 数据验证
|
||||||
|
node_ids = set(node['id'] for node in all_nodes)
|
||||||
|
validation_report = self.builder.validate_relationships(relationships, node_ids)
|
||||||
|
|
||||||
|
# 保存最终结果
|
||||||
|
self.logger.info("\n步骤3: 保存最终结果")
|
||||||
|
nodes_df = pd.DataFrame(all_nodes)
|
||||||
|
nodes_df.to_csv(nodes_file, index=False, encoding='utf-8-sig')
|
||||||
|
self.logger.info(f"最终节点已保存: {nodes_file} ({len(all_nodes)}行)")
|
||||||
|
|
||||||
|
rels_df = pd.DataFrame(relationships)
|
||||||
|
rels_df.to_csv(rels_file, index=False, encoding='utf-8-sig')
|
||||||
|
self.logger.info(f"最终关系已保存: {rels_file} ({len(relationships)}行)")
|
||||||
|
|
||||||
|
# 生成报告
|
||||||
|
self.logger.info("\n步骤4: 生成报告")
|
||||||
|
self.generate_report(
|
||||||
|
report_file,
|
||||||
|
projects_data,
|
||||||
|
all_nodes,
|
||||||
|
relationships,
|
||||||
|
validation_report,
|
||||||
|
start_time
|
||||||
|
)
|
||||||
|
|
||||||
|
end_time = datetime.now()
|
||||||
|
duration = (end_time - start_time).total_seconds()
|
||||||
|
|
||||||
|
# 返回统计信息
|
||||||
|
result_stats = {
|
||||||
|
'total_projects': len(projects_data),
|
||||||
|
'total_nodes': len(all_nodes),
|
||||||
|
'total_relationships': len(relationships),
|
||||||
|
'entity_nodes': len(entity_nodes),
|
||||||
|
'project_nodes': len(project_nodes),
|
||||||
|
'duration_seconds': duration,
|
||||||
|
'validation_report': validation_report
|
||||||
|
}
|
||||||
|
|
||||||
|
self.logger.info("\n"+"="*60)
|
||||||
|
self.logger.info(f"抽取完成!耗时: {duration:.2f}秒")
|
||||||
|
self.logger.info(f"项目数: {result_stats['total_projects']}")
|
||||||
|
self.logger.info(f"节点数: {result_stats['total_nodes']}")
|
||||||
|
self.logger.info(f"关系数: {result_stats['total_relationships']}")
|
||||||
|
self.logger.info("="*60)
|
||||||
|
|
||||||
|
return result_stats
|
||||||
|
|
||||||
|
def generate_report(
|
||||||
|
self,
|
||||||
|
report_file: Path,
|
||||||
|
projects_data: List[Dict[str, str]],
|
||||||
|
nodes: List[Dict[str, Any]],
|
||||||
|
relationships: List[Dict[str, Any]],
|
||||||
|
validation_report: Dict[str, Any],
|
||||||
|
start_time: datetime
|
||||||
|
):
|
||||||
|
"""
|
||||||
|
生成抽取报告
|
||||||
|
|
||||||
|
Args:
|
||||||
|
report_file: 报告文件路径
|
||||||
|
projects_data: 项目数据
|
||||||
|
nodes: 节点列表
|
||||||
|
relationships: 关系列表
|
||||||
|
validation_report: 验证报告
|
||||||
|
start_time: 开始时间
|
||||||
|
"""
|
||||||
|
report_lines = []
|
||||||
|
|
||||||
|
# 标题
|
||||||
|
report_lines.append("# 深度文化实体抽取报告\n")
|
||||||
|
report_lines.append(f"**生成时间**: {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n")
|
||||||
|
report_lines.append(f"**耗时**: {(datetime.now() - start_time).total_seconds():.2f}秒\n")
|
||||||
|
report_lines.append("---\n\n")
|
||||||
|
|
||||||
|
# 1. 处理统计
|
||||||
|
report_lines.append("## 1. 处理统计\n\n")
|
||||||
|
report_lines.append(f"- **处理项目数**: {len(projects_data)}\n")
|
||||||
|
report_lines.append(f"- **总节点数**: {len(nodes)}\n")
|
||||||
|
report_lines.append(f"- **总关系数**: {len(relationships)}\n")
|
||||||
|
report_lines.append(f"- **实体节点数**: {len([n for n in nodes if n['type'] != 'ICH_Project'])}\n")
|
||||||
|
report_lines.append(f"- **项目节点数**: {len([n for n in nodes if n['type'] == 'ICH_Project'])}\n")
|
||||||
|
|
||||||
|
# 2. 节点类型分布
|
||||||
|
report_lines.append("\n## 2. 节点类型分布\n\n")
|
||||||
|
node_types = {}
|
||||||
|
for node in nodes:
|
||||||
|
node_type = node['type']
|
||||||
|
node_types[node_type] = node_types.get(node_type, 0) + 1
|
||||||
|
|
||||||
|
report_lines.append("| 节点类型 | 数量 | 占比 |\n")
|
||||||
|
report_lines.append("|---------|------|------|\n")
|
||||||
|
for node_type, count in sorted(node_types.items()):
|
||||||
|
percentage = (count / len(nodes) * 100) if len(nodes) > 0 else 0
|
||||||
|
report_lines.append(f"| {node_type} | {count} | {percentage:.1f}% |\n")
|
||||||
|
|
||||||
|
# 3. 关系类型分布
|
||||||
|
report_lines.append("\n## 3. 关系类型分布\n\n")
|
||||||
|
rel_types = {}
|
||||||
|
for rel in relationships:
|
||||||
|
rel_type = rel['type']
|
||||||
|
rel_types[rel_type] = rel_types.get(rel_type, 0) + 1
|
||||||
|
|
||||||
|
report_lines.append("| 关系类型 | 数量 | 占比 |\n")
|
||||||
|
report_lines.append("|---------|------|------|\n")
|
||||||
|
for rel_type, count in sorted(rel_types.items()):
|
||||||
|
percentage = (count / len(relationships) * 100) if len(relationships) > 0 else 0
|
||||||
|
report_lines.append(f"| {rel_type} | {count} | {percentage:.1f}% |\n")
|
||||||
|
|
||||||
|
# 4. 数据质量
|
||||||
|
report_lines.append("\n## 4. 数据质量\n\n")
|
||||||
|
report_lines.append(f"- **有效关系**: {validation_report['valid_relationships']}\n")
|
||||||
|
report_lines.append(f"- **无效关系**: {validation_report['invalid_relationships']}\n")
|
||||||
|
|
||||||
|
if validation_report['broken_links']:
|
||||||
|
report_lines.append(f"\n**断链警告**: {len(validation_report['broken_links'])}个\n")
|
||||||
|
report_lines.append("```json\n")
|
||||||
|
report_lines.append(json.dumps(validation_report['broken_links'][:10], ensure_ascii=False, indent=2))
|
||||||
|
if len(validation_report['broken_links']) > 10:
|
||||||
|
report_lines.append(f"\n... (还有{len(validation_report['broken_links'])-10}个)")
|
||||||
|
report_lines.append("\n```\n")
|
||||||
|
|
||||||
|
# 5. 项目详情(前5个)
|
||||||
|
report_lines.append("\n## 5. 项目抽取详情(前5个)\n\n")
|
||||||
|
|
||||||
|
for idx, (project_data, node) in enumerate(zip(projects_data[:5], nodes[:5])):
|
||||||
|
if node['type'] != 'ICH_Project':
|
||||||
|
continue
|
||||||
|
|
||||||
|
report_lines.append(f"### {idx+1}. {project_data['project_name']}\n\n")
|
||||||
|
report_lines.append(f"**项目ID**: {project_data['project_id']}\n\n")
|
||||||
|
report_lines.append(f"**备注**: {project_data['remark_text'][:200]}...\n\n")
|
||||||
|
|
||||||
|
# 查找相关关系
|
||||||
|
related_rels = [r for r in relationships if r['source'] == project_data['project_id']]
|
||||||
|
if related_rels:
|
||||||
|
report_lines.append(f"**关系数**: {len(related_rels)}\n\n")
|
||||||
|
report_lines.append("| 关系类型 | 目标实体 |\n")
|
||||||
|
report_lines.append("|---------|---------|\n")
|
||||||
|
for rel in related_rels[:10]:
|
||||||
|
target_node = next((n for n in nodes if n['id'] == rel['target']), None)
|
||||||
|
if target_node:
|
||||||
|
report_lines.append(f"| {rel['type']} | {target_node['label']} |\n")
|
||||||
|
if len(related_rels) > 10:
|
||||||
|
report_lines.append(f"| ... | 还有{len(related_rels)-10}个关系 |\n")
|
||||||
|
|
||||||
|
report_lines.append("\n")
|
||||||
|
|
||||||
|
# 6. 输出文件
|
||||||
|
report_lines.append("## 6. 输出文件\n\n")
|
||||||
|
report_lines.append(f"- **节点文件**: `{report_file.parent / (report_file.stem.replace('_report', '') + '_nodes.csv')}`\n")
|
||||||
|
report_lines.append(f"- **关系文件**: `{report_file.parent / (report_file.stem.replace('_report', '') + '_rels.csv')}`\n")
|
||||||
|
|
||||||
|
# 保存报告
|
||||||
|
with open(report_file, 'w', encoding='utf-8') as f:
|
||||||
|
f.writelines(report_lines)
|
||||||
|
|
||||||
|
self.logger.info(f"报告已保存: {report_file}")
|
||||||
|
|
||||||
|
|
||||||
|
async def main():
|
||||||
|
"""主函数"""
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 配置文件
|
||||||
|
config_file = r"E:\Project\2026_KG_ICH\dofile\kg_project\config\deep_extraction_config.yaml"
|
||||||
|
|
||||||
|
# 数据文件:从命令行参数获取,或使用默认完整数据文件
|
||||||
|
if len(sys.argv) > 1:
|
||||||
|
excel_file = sys.argv[1]
|
||||||
|
else:
|
||||||
|
excel_file = r"E:\Project\2026_KG_ICH\data\黑龙江国家级和省级非遗名单.xlsx"
|
||||||
|
|
||||||
|
# 创建流程
|
||||||
|
pipeline = DeepExtractionPipeline(config_file)
|
||||||
|
|
||||||
|
# 执行抽取
|
||||||
|
results = await pipeline.run_extraction(
|
||||||
|
excel_file=excel_file,
|
||||||
|
output_prefix="", # 不使用前缀(完整抽取)
|
||||||
|
max_rows=None # 读取所有行
|
||||||
|
)
|
||||||
|
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("抽取完成!")
|
||||||
|
print("="*60)
|
||||||
|
print(f"\n结果统计:")
|
||||||
|
print(json.dumps(results, indent=2, ensure_ascii=False))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
asyncio.run(main())
|
||||||
@@ -0,0 +1,420 @@
|
|||||||
|
# -*- coding: utf-8 -*-
|
||||||
|
"""
|
||||||
|
深度文化实体抽取器
|
||||||
|
利用LLM从非遗项目备注中抽取深层文化实体和关系
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import asyncio
|
||||||
|
import logging
|
||||||
|
from typing import Dict, List, Any, Optional
|
||||||
|
from pathlib import Path
|
||||||
|
import yaml
|
||||||
|
from langchain_deepseek import ChatDeepSeek
|
||||||
|
from langchain_core.messages import HumanMessage, SystemMessage
|
||||||
|
|
||||||
|
|
||||||
|
class DeepEntityExtractor:
|
||||||
|
"""深度文化实体抽取器"""
|
||||||
|
|
||||||
|
def __init__(self, config_file: str):
|
||||||
|
"""
|
||||||
|
初始化抽取器
|
||||||
|
|
||||||
|
Args:
|
||||||
|
config_file: 配置文件路径
|
||||||
|
"""
|
||||||
|
# 加载配置
|
||||||
|
with open(config_file, 'r', encoding='utf-8') as f:
|
||||||
|
self.config = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 设置日志
|
||||||
|
log_config = self.config.get('logging', {})
|
||||||
|
log_file = log_config.get('log_file', 'logs/deep_extraction.log')
|
||||||
|
Path(log_file).parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# 创建logger
|
||||||
|
self.logger = logging.getLogger(__name__)
|
||||||
|
self.logger.setLevel(getattr(logging, log_config.get('level', 'INFO')))
|
||||||
|
|
||||||
|
# 清除已有的handlers
|
||||||
|
self.logger.handlers.clear()
|
||||||
|
|
||||||
|
# 文件handler(实时刷新)
|
||||||
|
file_handler = logging.FileHandler(log_file, encoding='utf-8')
|
||||||
|
file_handler.setLevel(getattr(logging, log_config.get('level', 'INFO')))
|
||||||
|
file_formatter = logging.Formatter(log_config.get('format', '%(asctime)s - %(name)s - %(levelname)s - %(message)s'))
|
||||||
|
file_handler.setFormatter(file_formatter)
|
||||||
|
# 强制实时刷新
|
||||||
|
file_handler.flush = lambda: file_handler.stream.flush()
|
||||||
|
self.logger.addHandler(file_handler)
|
||||||
|
|
||||||
|
# 控制台handler
|
||||||
|
console_handler = logging.StreamHandler()
|
||||||
|
console_handler.setLevel(getattr(logging, log_config.get('level', 'INFO')))
|
||||||
|
console_formatter = logging.Formatter(log_config.get('format', '%(asctime)s - %(name)s - %(levelname)s - %(message)s'))
|
||||||
|
console_handler.setFormatter(console_formatter)
|
||||||
|
self.logger.addHandler(console_handler)
|
||||||
|
|
||||||
|
# 加载实体本体配置
|
||||||
|
ontology_file = Path(config_file).parent / 'entity_ontology.yaml'
|
||||||
|
with open(ontology_file, 'r', encoding='utf-8') as f:
|
||||||
|
self.ontology = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 初始化LLM
|
||||||
|
self._init_llm()
|
||||||
|
|
||||||
|
# 构建提示词
|
||||||
|
self.system_prompt = self._build_system_prompt()
|
||||||
|
|
||||||
|
self.logger.info("DeepEntityExtractor初始化完成")
|
||||||
|
|
||||||
|
def _init_llm(self):
|
||||||
|
"""初始化LLM模型"""
|
||||||
|
llm_config = self.config['llm']
|
||||||
|
|
||||||
|
# 读取API密钥(从现有的api_keys.yaml)
|
||||||
|
api_keys_file = Path(__file__).parent.parent.parent / 'config' / 'api_keys.yaml'
|
||||||
|
with open(api_keys_file, 'r', encoding='utf-8') as f:
|
||||||
|
api_keys = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 适配API密钥格式
|
||||||
|
api_key = api_keys.get('deepseek_api_key', '')
|
||||||
|
if not api_key:
|
||||||
|
# 尝试嵌套格式
|
||||||
|
api_key = api_keys.get('deepseek', {}).get('api_key', '')
|
||||||
|
|
||||||
|
self.llm = ChatDeepSeek(
|
||||||
|
model=llm_config['model'],
|
||||||
|
temperature=llm_config['temperature'],
|
||||||
|
max_tokens=llm_config['max_tokens'],
|
||||||
|
api_key=api_key
|
||||||
|
)
|
||||||
|
|
||||||
|
self.logger.info(f"LLM初始化完成: {llm_config['model']}")
|
||||||
|
|
||||||
|
def _build_system_prompt(self) -> str:
|
||||||
|
"""构建系统提示词"""
|
||||||
|
entity_prompts = self.config.get('entity_type_prompts', {})
|
||||||
|
relationship_prompts = self.config.get('relationship_type_prompts', {})
|
||||||
|
|
||||||
|
prompt = """你是一位非物质文化遗产领域的专家,擅长从项目描述中识别深层文化实体和关系。
|
||||||
|
|
||||||
|
【任务】
|
||||||
|
请从非遗项目描述中识别以下9类深层文化实体和9种深层关系:
|
||||||
|
|
||||||
|
=== 实体类型 ===
|
||||||
|
"""
|
||||||
|
|
||||||
|
# 添加实体类型说明
|
||||||
|
for entity_type, description in entity_prompts.items():
|
||||||
|
prompt += f"{entity_type}:{description}\n"
|
||||||
|
|
||||||
|
prompt += "\n=== 关系类型 ===\n"
|
||||||
|
|
||||||
|
# 添加关系类型说明
|
||||||
|
for rel_type, description in relationship_prompts.items():
|
||||||
|
prompt += f"{rel_type}:{description}\n"
|
||||||
|
|
||||||
|
prompt += """
|
||||||
|
【输出要求】
|
||||||
|
1. 严格按JSON格式输出,不要包含任何其他文本
|
||||||
|
2. 只输出明确的、文本中提到的实体和关系,不要臆测
|
||||||
|
3. 实体text必须从原文中提取,不要自行改写
|
||||||
|
4. 同一个实体只识别一次,避免重复
|
||||||
|
5. 关系必须基于文本中的明确描述
|
||||||
|
6. 如果某类实体或关系不存在,相应数组为空
|
||||||
|
7. 确保JSON格式正确,可以被Python解析
|
||||||
|
|
||||||
|
【JSON输出格式】
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"entities": [
|
||||||
|
{
|
||||||
|
"text": "实体文本(如:鱼皮)",
|
||||||
|
"type": "实体类型(如:Material)",
|
||||||
|
"attributes": {
|
||||||
|
"name": "标准名称",
|
||||||
|
"category": "类别(如:动物材料)",
|
||||||
|
"description": "详细描述"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"relationships": [
|
||||||
|
{
|
||||||
|
"source_entity": "ICH-{项目ID}",
|
||||||
|
"target_entity": "目标实体文本",
|
||||||
|
"type": "关系类型",
|
||||||
|
"properties": {
|
||||||
|
"description": "关系描述",
|
||||||
|
"context": "上下文信息"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"total_entities": 0,
|
||||||
|
"total_relationships": 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
"""
|
||||||
|
|
||||||
|
return prompt
|
||||||
|
|
||||||
|
async def extract_from_remark(
|
||||||
|
self,
|
||||||
|
project_id: str,
|
||||||
|
project_name: str,
|
||||||
|
remark_text: str
|
||||||
|
) -> Optional[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
从单条备注抽取实体和关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
project_id: 项目ID(如ICH-1)
|
||||||
|
project_name: 项目名称
|
||||||
|
remark_text: 备注文本
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
抽取结果(包含entities, relationships, summary)
|
||||||
|
"""
|
||||||
|
if not remark_text or len(remark_text.strip()) < 10:
|
||||||
|
self.logger.warning(f"项目 {project_id} 备注文本过短,跳过抽取")
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 构建用户提示词
|
||||||
|
user_prompt = f"""【项目信息】
|
||||||
|
项目名称:{project_name}
|
||||||
|
项目ID:{project_id}
|
||||||
|
|
||||||
|
【描述文本】
|
||||||
|
{remark_text}
|
||||||
|
|
||||||
|
请从上述描述中识别深层文化实体和关系,严格按JSON格式输出。"""
|
||||||
|
|
||||||
|
try:
|
||||||
|
# 调用LLM
|
||||||
|
messages = [
|
||||||
|
SystemMessage(content=self.system_prompt),
|
||||||
|
HumanMessage(content=user_prompt)
|
||||||
|
]
|
||||||
|
|
||||||
|
max_retries = self.config['llm']['max_retries']
|
||||||
|
retry_delay = self.config['llm']['retry_delay']
|
||||||
|
request_timeout = self.config['llm'].get('request_timeout', 120)
|
||||||
|
|
||||||
|
for attempt in range(max_retries):
|
||||||
|
try:
|
||||||
|
self.logger.info(f"项目 {project_id} 开始LLM调用(尝试{attempt+1}/{max_retries})")
|
||||||
|
|
||||||
|
# 添加超时控制
|
||||||
|
response = await asyncio.wait_for(
|
||||||
|
self.llm.ainvoke(messages),
|
||||||
|
timeout=request_timeout
|
||||||
|
)
|
||||||
|
result_text = response.content
|
||||||
|
|
||||||
|
self.logger.info(f"项目 {project_id} LLM调用成功,开始提取JSON")
|
||||||
|
|
||||||
|
# 提取JSON
|
||||||
|
extraction_result = self._extract_json(result_text)
|
||||||
|
|
||||||
|
if extraction_result:
|
||||||
|
# 验证结果
|
||||||
|
if self._validate_extraction(extraction_result):
|
||||||
|
self.logger.info(f"项目 {project_id} 抽取成功:{extraction_result['summary']['total_entities']}个实体,{extraction_result['summary']['total_relationships']}个关系")
|
||||||
|
return extraction_result
|
||||||
|
else:
|
||||||
|
self.logger.warning(f"项目 {project_id} 抽取结果验证失败")
|
||||||
|
else:
|
||||||
|
self.logger.warning(f"项目 {project_id} JSON提取失败")
|
||||||
|
|
||||||
|
except asyncio.TimeoutError:
|
||||||
|
self.logger.error(f"项目 {project_id} LLM调用超时({request_timeout}秒)(尝试{attempt+1}/{max_retries})")
|
||||||
|
if attempt < max_retries - 1:
|
||||||
|
await asyncio.sleep(retry_delay)
|
||||||
|
else:
|
||||||
|
self.logger.error(f"项目 {project_id} 达到最大重试次数,放弃抽取")
|
||||||
|
raise
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"项目 {project_id} LLM调用失败(尝试{attempt+1}/{max_retries}): {type(e).__name__}: {str(e)}")
|
||||||
|
if attempt < max_retries - 1:
|
||||||
|
await asyncio.sleep(retry_delay)
|
||||||
|
else:
|
||||||
|
raise
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"项目 {project_id} 抽取失败: {str(e)}", exc_info=True)
|
||||||
|
return None
|
||||||
|
|
||||||
|
def _extract_json(self, text: str) -> Optional[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
从文本中提取JSON
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text: LLM返回的文本
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
解析后的JSON对象,失败返回None
|
||||||
|
"""
|
||||||
|
# 尝试直接解析
|
||||||
|
try:
|
||||||
|
return json.loads(text)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
# 尝试提取JSON块
|
||||||
|
import re
|
||||||
|
json_pattern = r'```json\s*(.*?)\s*```'
|
||||||
|
match = re.search(json_pattern, text, re.DOTALL)
|
||||||
|
|
||||||
|
if match:
|
||||||
|
try:
|
||||||
|
return json.loads(match.group(1))
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
# 尝试提取花括号内容
|
||||||
|
brace_pattern = r'\{.*\}'
|
||||||
|
match = re.search(brace_pattern, text, re.DOTALL)
|
||||||
|
|
||||||
|
if match:
|
||||||
|
try:
|
||||||
|
return json.loads(match.group(0))
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
pass
|
||||||
|
|
||||||
|
self.logger.error(f"无法从文本中提取有效JSON: {text[:200]}...")
|
||||||
|
return None
|
||||||
|
|
||||||
|
def _validate_extraction(self, result: Dict[str, Any]) -> bool:
|
||||||
|
"""
|
||||||
|
验证抽取结果
|
||||||
|
|
||||||
|
Args:
|
||||||
|
result: 抽取结果
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
是否有效
|
||||||
|
"""
|
||||||
|
# 检查必需字段
|
||||||
|
required_fields = ['entities', 'relationships', 'summary']
|
||||||
|
for field in required_fields:
|
||||||
|
if field not in result:
|
||||||
|
self.logger.error(f"缺少必需字段: {field}")
|
||||||
|
return False
|
||||||
|
|
||||||
|
# 检查entities格式
|
||||||
|
if not isinstance(result['entities'], list):
|
||||||
|
self.logger.error("entities必须是列表")
|
||||||
|
return False
|
||||||
|
|
||||||
|
for entity in result['entities']:
|
||||||
|
if not all(k in entity for k in ['text', 'type', 'attributes']):
|
||||||
|
self.logger.error(f"实体缺少必需字段: {entity}")
|
||||||
|
return False
|
||||||
|
|
||||||
|
# 检查实体类型是否合法
|
||||||
|
entity_type = entity['type']
|
||||||
|
if entity_type not in self.ontology['entity_types']:
|
||||||
|
self.logger.warning(f"未知实体类型: {entity_type}")
|
||||||
|
|
||||||
|
# 检查relationships格式
|
||||||
|
if not isinstance(result['relationships'], list):
|
||||||
|
self.logger.error("relationships必须是列表")
|
||||||
|
return False
|
||||||
|
|
||||||
|
for rel in result['relationships']:
|
||||||
|
if not all(k in rel for k in ['source_entity', 'target_entity', 'type', 'properties']):
|
||||||
|
self.logger.error(f"关系缺少必需字段: {rel}")
|
||||||
|
return False
|
||||||
|
|
||||||
|
# 检查关系类型是否合法
|
||||||
|
rel_type = rel['type']
|
||||||
|
if rel_type not in self.ontology['relationship_types']:
|
||||||
|
self.logger.warning(f"未知关系类型: {rel_type}")
|
||||||
|
|
||||||
|
return True
|
||||||
|
|
||||||
|
async def batch_extract(
|
||||||
|
self,
|
||||||
|
projects_data: List[Dict[str, str]],
|
||||||
|
batch_size: int = None,
|
||||||
|
progress_callback=None
|
||||||
|
) -> List[Optional[Dict[str, Any]]]:
|
||||||
|
"""
|
||||||
|
批量抽取实体和关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
projects_data: 项目数据列表,每项包含project_id, project_name, remark_text
|
||||||
|
batch_size: 批次大小(从配置文件读取)
|
||||||
|
progress_callback: 进度回调函数,每批次完成后调用
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
抽取结果列表
|
||||||
|
"""
|
||||||
|
if batch_size is None:
|
||||||
|
batch_size = self.config['batch_processing']['batch_size']
|
||||||
|
|
||||||
|
batch_timeout = self.config['llm'].get('batch_timeout', 300)
|
||||||
|
incremental_save = self.config['batch_processing'].get('incremental_save', False)
|
||||||
|
|
||||||
|
results = []
|
||||||
|
total = len(projects_data)
|
||||||
|
|
||||||
|
self.logger.info(f"="*60)
|
||||||
|
self.logger.info(f"开始批量抽取,共{total}个项目,批次大小{batch_size}")
|
||||||
|
self.logger.info(f"增量保存: {'启用' if incremental_save else '禁用'}")
|
||||||
|
self.logger.info(f"请求超时: {self.config['llm'].get('request_timeout', 120)}秒")
|
||||||
|
self.logger.info(f"批次超时: {batch_timeout}秒")
|
||||||
|
self.logger.info(f"="*60)
|
||||||
|
|
||||||
|
for i in range(0, total, batch_size):
|
||||||
|
batch = projects_data[i:i+batch_size]
|
||||||
|
batch_num = i // batch_size + 1
|
||||||
|
total_batches = (total + batch_size - 1) // batch_size
|
||||||
|
|
||||||
|
self.logger.info(f"-"*60)
|
||||||
|
self.logger.info(f"开始处理批次 {batch_num}/{total_batches},包含{len(batch)}个项目")
|
||||||
|
for item in batch:
|
||||||
|
self.logger.info(f" - {item['project_id']}: {item['project_name']}")
|
||||||
|
|
||||||
|
# 并发处理当前批次
|
||||||
|
batch_tasks = [
|
||||||
|
self.extract_from_remark(
|
||||||
|
item['project_id'],
|
||||||
|
item['project_name'],
|
||||||
|
item['remark_text']
|
||||||
|
)
|
||||||
|
for item in batch
|
||||||
|
]
|
||||||
|
|
||||||
|
try:
|
||||||
|
# 添加批次级别的超时控制
|
||||||
|
batch_results = await asyncio.wait_for(
|
||||||
|
asyncio.gather(*batch_tasks, return_exceptions=True),
|
||||||
|
timeout=batch_timeout
|
||||||
|
)
|
||||||
|
results.extend(batch_results)
|
||||||
|
|
||||||
|
# 统计本批次结果
|
||||||
|
success_count = sum(1 for r in batch_results if r is not None and not isinstance(r, Exception))
|
||||||
|
error_count = len(batch_results) - success_count
|
||||||
|
|
||||||
|
self.logger.info(f"批次 {batch_num}/{total_batches} 完成 - 成功: {success_count}, 失败: {error_count}")
|
||||||
|
|
||||||
|
# 调用进度回调(用于增量保存)
|
||||||
|
if progress_callback:
|
||||||
|
await progress_callback(batch_num, total_batches, results, projects_data[:i+len(batch)])
|
||||||
|
|
||||||
|
except asyncio.TimeoutError:
|
||||||
|
self.logger.error(f"批次 {batch_num}/{total_batches} 超时({batch_timeout}秒),部分请求失败")
|
||||||
|
# 将未完成的结果标记为None
|
||||||
|
for item in batch[len(results):]:
|
||||||
|
results.append(None)
|
||||||
|
|
||||||
|
self.logger.info(f"="*60)
|
||||||
|
self.logger.info(f"批量抽取完成,共处理{len(results)}个项目")
|
||||||
|
self.logger.info(f"="*60)
|
||||||
|
|
||||||
|
return results
|
||||||
@@ -0,0 +1,486 @@
|
|||||||
|
"""
|
||||||
|
DeepSeek实体识别模块
|
||||||
|
使用DeepSeek API进行非遗知识抽取
|
||||||
|
"""
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Dict, List, Optional, Any
|
||||||
|
import logging
|
||||||
|
from datetime import datetime
|
||||||
|
|
||||||
|
try:
|
||||||
|
from langchain_deepseek import ChatDeepSeek
|
||||||
|
from langchain_core.messages import HumanMessage, SystemMessage
|
||||||
|
except ImportError:
|
||||||
|
print("错误: 请先安装依赖包")
|
||||||
|
print("运行: pip install langchain-deepseek langchain-core")
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
|
class DeepSeekExtractor:
|
||||||
|
"""基于DeepSeek的知识抽取器"""
|
||||||
|
|
||||||
|
def __init__(self, config_file: str = 'config/ich_config.yaml'):
|
||||||
|
"""
|
||||||
|
初始化抽取器
|
||||||
|
|
||||||
|
Args:
|
||||||
|
config_file: 配置文件路径
|
||||||
|
"""
|
||||||
|
# 加载配置
|
||||||
|
self.config = self._load_config(config_file)
|
||||||
|
self.api_key = self._load_api_key()
|
||||||
|
|
||||||
|
# 初始化模型
|
||||||
|
self.model = ChatDeepSeek(
|
||||||
|
model=self.config['api']['model'],
|
||||||
|
api_key=self.api_key,
|
||||||
|
temperature=self.config['api']['temperature'],
|
||||||
|
max_tokens=self.config['api']['max_tokens']
|
||||||
|
)
|
||||||
|
|
||||||
|
# 设置日志
|
||||||
|
self.logger = self._setup_logger()
|
||||||
|
|
||||||
|
# 加载本体配置
|
||||||
|
self.entity_types = self.config['extraction']['entity_types']
|
||||||
|
self.relation_types = self.config['extraction']['relation_types']
|
||||||
|
|
||||||
|
def _setup_logger(self):
|
||||||
|
"""设置日志"""
|
||||||
|
log_dir = Path(self.config['paths']['logs_dir'])
|
||||||
|
log_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
log_file = log_dir / f"extraction_{datetime.now().strftime('%Y%m%d')}.log"
|
||||||
|
logging.basicConfig(
|
||||||
|
level=getattr(logging, self.config['logging']['level']),
|
||||||
|
format=self.config['logging']['format'],
|
||||||
|
handlers=[
|
||||||
|
logging.FileHandler(log_file, encoding='utf-8'),
|
||||||
|
logging.StreamHandler()
|
||||||
|
]
|
||||||
|
)
|
||||||
|
return logging.getLogger(__name__)
|
||||||
|
|
||||||
|
def _load_config(self, config_file: str) -> Dict:
|
||||||
|
"""加载配置"""
|
||||||
|
config_path = Path(config_file)
|
||||||
|
if not config_path.exists():
|
||||||
|
raise FileNotFoundError(f"配置文件不存在: {config_file}")
|
||||||
|
|
||||||
|
with open(config_path, 'r', encoding='utf-8') as f:
|
||||||
|
return yaml.safe_load(f)
|
||||||
|
|
||||||
|
def _load_api_key(self) -> str:
|
||||||
|
"""加载API密钥"""
|
||||||
|
api_config_file = Path(self.config['api']['config_file'])
|
||||||
|
if not api_config_file.exists():
|
||||||
|
raise FileNotFoundError(f"API配置文件不存在: {api_config_file}")
|
||||||
|
|
||||||
|
with open(api_config_file, 'r', encoding='utf-8') as f:
|
||||||
|
api_config = yaml.safe_load(f)
|
||||||
|
|
||||||
|
return api_config['deepseek_api_key']
|
||||||
|
|
||||||
|
def extract_json_from_response(self, response_text: str) -> str:
|
||||||
|
"""
|
||||||
|
从API响应中提取JSON内容
|
||||||
|
|
||||||
|
Args:
|
||||||
|
response_text: API响应文本
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
str: 提取的JSON字符串
|
||||||
|
"""
|
||||||
|
# 如果响应包含markdown代码块,提取其中的JSON
|
||||||
|
if '```json' in response_text:
|
||||||
|
start = response_text.find('```json') + 7
|
||||||
|
end = response_text.find('```', start)
|
||||||
|
if start > 6 and end > start:
|
||||||
|
return response_text[start:end].strip()
|
||||||
|
elif '```' in response_text:
|
||||||
|
start = response_text.find('```') + 3
|
||||||
|
end = response_text.find('```', start)
|
||||||
|
if start > 2 and end > start:
|
||||||
|
content = response_text[start:end].strip()
|
||||||
|
if not content.startswith('json'):
|
||||||
|
return content
|
||||||
|
return content[5:].strip() if content.startswith('json') else content.strip()
|
||||||
|
|
||||||
|
# 否则直接返回原始文本
|
||||||
|
return response_text.strip()
|
||||||
|
|
||||||
|
async def extract_entities_from_text(
|
||||||
|
self,
|
||||||
|
text: str,
|
||||||
|
project_name: str = "",
|
||||||
|
max_retries: int = 3
|
||||||
|
) -> Optional[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
从文本中抽取实体
|
||||||
|
|
||||||
|
Args:
|
||||||
|
text: 输入文本
|
||||||
|
project_name: 项目名称(可选)
|
||||||
|
max_retries: 最大重试次数
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Dict: 抽取结果
|
||||||
|
"""
|
||||||
|
# 构建提示词
|
||||||
|
entity_types_str = "、".join(self.entity_types)
|
||||||
|
|
||||||
|
prompt = f"""
|
||||||
|
你是一位非物质文化遗产领域的专家。请从以下文本中识别非遗相关实体,并提取属性。
|
||||||
|
|
||||||
|
项目名称:{project_name}
|
||||||
|
|
||||||
|
文本内容:
|
||||||
|
{text}
|
||||||
|
|
||||||
|
请识别以下类型的实体:
|
||||||
|
{entity_types_str}
|
||||||
|
|
||||||
|
对于每个实体,请提取以下信息:
|
||||||
|
1. 实体文本
|
||||||
|
2. 实体类型
|
||||||
|
3. 相关属性(如:民族、地点、时间、技艺特点等)
|
||||||
|
|
||||||
|
输出格式(JSON):
|
||||||
|
{{
|
||||||
|
"entities": [
|
||||||
|
{{
|
||||||
|
"text": "实体文本",
|
||||||
|
"type": "实体类型",
|
||||||
|
"attributes": {{
|
||||||
|
"ethnic_group": "民族(如果适用)",
|
||||||
|
"location": "地点(如果适用)",
|
||||||
|
"time_period": "时期(如果适用)",
|
||||||
|
"skill_feature": "技艺特点(如果适用)",
|
||||||
|
"cultural_value": "文化价值(如果适用)"
|
||||||
|
}}
|
||||||
|
}}
|
||||||
|
],
|
||||||
|
"relationships": [
|
||||||
|
{{
|
||||||
|
"from": "实体1",
|
||||||
|
"to": "实体2",
|
||||||
|
"type": "关系类型",
|
||||||
|
"description": "关系描述"
|
||||||
|
}}
|
||||||
|
]
|
||||||
|
}}
|
||||||
|
|
||||||
|
请确保输出是有效的JSON格式。
|
||||||
|
"""
|
||||||
|
|
||||||
|
# 调用API
|
||||||
|
for attempt in range(max_retries):
|
||||||
|
try:
|
||||||
|
self.logger.info(f"开始抽取实体 (尝试 {attempt + 1}/{max_retries})")
|
||||||
|
|
||||||
|
messages = [HumanMessage(content=prompt)]
|
||||||
|
response = await self.model.ainvoke(messages)
|
||||||
|
result_text = response.content
|
||||||
|
|
||||||
|
# 提取JSON
|
||||||
|
json_text = self.extract_json_from_response(result_text)
|
||||||
|
|
||||||
|
# 解析JSON
|
||||||
|
try:
|
||||||
|
result = json.loads(json_text)
|
||||||
|
|
||||||
|
# 添加元数据
|
||||||
|
result['metadata'] = {
|
||||||
|
'project_name': project_name,
|
||||||
|
'extraction_time': datetime.now().isoformat(),
|
||||||
|
'text_length': len(text),
|
||||||
|
'model': self.config['api']['model']
|
||||||
|
}
|
||||||
|
|
||||||
|
self.logger.info(f"成功抽取 {len(result.get('entities', []))} 个实体")
|
||||||
|
return result
|
||||||
|
|
||||||
|
except json.JSONDecodeError as e:
|
||||||
|
self.logger.warning(f"JSON解析失败: {str(e)}")
|
||||||
|
self.logger.debug(f"响应文本: {result_text[:500]}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
raise
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"API调用失败 (尝试 {attempt + 1}/{max_retries}): {str(e)}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
return None
|
||||||
|
|
||||||
|
# 等待后重试
|
||||||
|
await asyncio.sleep(self.config['processing']['retry_delay'])
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
async def extract_relationships(
|
||||||
|
self,
|
||||||
|
entities: List[Dict],
|
||||||
|
context: str = "",
|
||||||
|
max_retries: int = 3
|
||||||
|
) -> Optional[List[Dict[str, Any]]]:
|
||||||
|
"""
|
||||||
|
抽取实体间的关系
|
||||||
|
|
||||||
|
Args:
|
||||||
|
entities: 实体列表
|
||||||
|
context: 上下文文本
|
||||||
|
max_retries: 最大重试次数
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List[Dict]: 关系列表
|
||||||
|
"""
|
||||||
|
if not entities:
|
||||||
|
return []
|
||||||
|
|
||||||
|
# 构建实体描述
|
||||||
|
entity_descriptions = []
|
||||||
|
for entity in entities:
|
||||||
|
desc = f"- {entity.get('text', '')} ({entity.get('type', '')})"
|
||||||
|
if entity.get('attributes'):
|
||||||
|
attrs = ", ".join([f"{k}={v}" for k, v in entity['attributes'].items() if v])
|
||||||
|
if attrs:
|
||||||
|
desc += f" [{attrs}]"
|
||||||
|
entity_descriptions.append(desc)
|
||||||
|
|
||||||
|
entities_str = "\n".join(entity_descriptions)
|
||||||
|
relations_str = "、".join(self.relation_types)
|
||||||
|
|
||||||
|
prompt = f"""
|
||||||
|
基于以下实体和上下文,识别实体间的关系:
|
||||||
|
|
||||||
|
实体列表:
|
||||||
|
{entities_str}
|
||||||
|
|
||||||
|
上下文:
|
||||||
|
{context}
|
||||||
|
|
||||||
|
可能的关系类型:
|
||||||
|
{relations_str}
|
||||||
|
|
||||||
|
对于每个关系,请提供:
|
||||||
|
1. 头实体(from)
|
||||||
|
2. 尾实体(to)
|
||||||
|
3. 关系类型
|
||||||
|
4. 关系描述
|
||||||
|
|
||||||
|
输出格式(JSON):
|
||||||
|
{{
|
||||||
|
"relationships": [
|
||||||
|
{{
|
||||||
|
"from": "实体1文本",
|
||||||
|
"to": "实体2文本",
|
||||||
|
"type": "关系类型",
|
||||||
|
"description": "关系描述",
|
||||||
|
"confidence": 0.9
|
||||||
|
}}
|
||||||
|
]
|
||||||
|
}}
|
||||||
|
|
||||||
|
请确保输出是有效的JSON格式。
|
||||||
|
"""
|
||||||
|
|
||||||
|
# 调用API
|
||||||
|
for attempt in range(max_retries):
|
||||||
|
try:
|
||||||
|
messages = [HumanMessage(content=prompt)]
|
||||||
|
response = await self.model.ainvoke(messages)
|
||||||
|
result_text = response.content
|
||||||
|
|
||||||
|
# 提取JSON
|
||||||
|
json_text = self.extract_json_from_response(result_text)
|
||||||
|
|
||||||
|
# 解析JSON
|
||||||
|
try:
|
||||||
|
result = json.loads(json_text)
|
||||||
|
relationships = result.get('relationships', [])
|
||||||
|
|
||||||
|
self.logger.info(f"成功抽取 {len(relationships)} 个关系")
|
||||||
|
return relationships
|
||||||
|
|
||||||
|
except json.JSONDecodeError as e:
|
||||||
|
self.logger.warning(f"JSON解析失败: {str(e)}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
raise
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"API调用失败 (尝试 {attempt + 1}/{max_retries}): {str(e)}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
return None
|
||||||
|
|
||||||
|
await asyncio.sleep(self.config['processing']['retry_delay'])
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
async def enrich_project_entity(
|
||||||
|
self,
|
||||||
|
project_data: Dict[str, Any],
|
||||||
|
max_retries: int = 3
|
||||||
|
) -> Optional[Dict[str, Any]]:
|
||||||
|
"""
|
||||||
|
丰富非遗项目实体信息
|
||||||
|
|
||||||
|
Args:
|
||||||
|
project_data: 项目数据
|
||||||
|
max_retries: 最大重试次数
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
Dict: 丰富后的实体信息
|
||||||
|
"""
|
||||||
|
project_name = project_data.get('properties', {}).get('name', '')
|
||||||
|
description = project_data.get('properties', {}).get('description', '')
|
||||||
|
|
||||||
|
prompt = f"""
|
||||||
|
你是一位非物质文化遗产领域的专家。请分析以下非遗项目,提取和丰富实体信息。
|
||||||
|
|
||||||
|
项目名称:{project_name}
|
||||||
|
|
||||||
|
项目描述:
|
||||||
|
{description}
|
||||||
|
|
||||||
|
请提取以下信息:
|
||||||
|
|
||||||
|
1. **民族特色**:是否与特定少数民族相关?(满族、赫哲族、鄂伦春族等)
|
||||||
|
2. **地域特征**:体现哪些黑龙江地域特征?(寒地、冰雪、森林、江河等)
|
||||||
|
3. **技艺特点**:核心技艺特点是什么?
|
||||||
|
4. **文化价值**:有哪些重要的文化价值?
|
||||||
|
5. **传承方式**:如何传承?(家族传承、师徒制度、口传心授等)
|
||||||
|
6. **濒危状况**:是否濒危?原因是什么?
|
||||||
|
7. **保护措施**:有哪些保护措施?
|
||||||
|
|
||||||
|
输出格式(JSON):
|
||||||
|
{{
|
||||||
|
"ethnic_features": ["民族1", "民族2"],
|
||||||
|
"regional_characteristics": ["特征1", "特征2"],
|
||||||
|
"skill_features": ["技艺特点1", "技艺特点2"],
|
||||||
|
"cultural_values": ["价值1", "价值2"],
|
||||||
|
"transmission_methods": ["方式1", "方式2"],
|
||||||
|
"endangerment_status": "濒危状况描述",
|
||||||
|
"protection_measures": ["措施1", "措施2"],
|
||||||
|
"summary": "项目总结(50字以内)"
|
||||||
|
}}
|
||||||
|
|
||||||
|
请确保输出是有效的JSON格式。
|
||||||
|
"""
|
||||||
|
|
||||||
|
# 调用API
|
||||||
|
for attempt in range(max_retries):
|
||||||
|
try:
|
||||||
|
messages = [HumanMessage(content=prompt)]
|
||||||
|
response = await self.model.ainvoke(messages)
|
||||||
|
result_text = response.content
|
||||||
|
|
||||||
|
# 提取JSON
|
||||||
|
json_text = self.extract_json_from_response(result_text)
|
||||||
|
|
||||||
|
# 解析JSON
|
||||||
|
try:
|
||||||
|
result = json.loads(json_text)
|
||||||
|
|
||||||
|
self.logger.info(f"成功丰富项目实体信息: {project_name}")
|
||||||
|
return result
|
||||||
|
|
||||||
|
except json.JSONDecodeError as e:
|
||||||
|
self.logger.warning(f"JSON解析失败: {str(e)}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
raise
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"API调用失败 (尝试 {attempt + 1}/{max_retries}): {str(e)}")
|
||||||
|
if attempt == max_retries - 1:
|
||||||
|
return None
|
||||||
|
|
||||||
|
await asyncio.sleep(self.config['processing']['retry_delay'])
|
||||||
|
|
||||||
|
return None
|
||||||
|
|
||||||
|
async def batch_extract(
|
||||||
|
self,
|
||||||
|
items: List[Dict[str, Any]],
|
||||||
|
concurrent: int = 5
|
||||||
|
) -> List[Optional[Dict[str, Any]]]:
|
||||||
|
"""
|
||||||
|
批量抽取
|
||||||
|
|
||||||
|
Args:
|
||||||
|
items: 待处理项目列表
|
||||||
|
concurrent: 并发数
|
||||||
|
|
||||||
|
Returns:
|
||||||
|
List[Dict]: 抽取结果列表
|
||||||
|
"""
|
||||||
|
results = []
|
||||||
|
batch_size = concurrent
|
||||||
|
|
||||||
|
for i in range(0, len(items), batch_size):
|
||||||
|
batch = items[i:i + batch_size]
|
||||||
|
self.logger.info(f"处理批次 {i//batch_size + 1}: {len(batch)} 个项目")
|
||||||
|
|
||||||
|
# 并发处理
|
||||||
|
tasks = []
|
||||||
|
for item in batch:
|
||||||
|
project_name = item.get('properties', {}).get('name', '')
|
||||||
|
description = item.get('properties', {}).get('description', '')
|
||||||
|
|
||||||
|
if description:
|
||||||
|
task = self.enrich_project_entity(item)
|
||||||
|
tasks.append(task)
|
||||||
|
else:
|
||||||
|
tasks.append(asyncio.sleep(0)) # 占位任务
|
||||||
|
|
||||||
|
batch_results = await asyncio.gather(*tasks, return_exceptions=True)
|
||||||
|
results.extend(batch_results)
|
||||||
|
|
||||||
|
# 显示进度
|
||||||
|
completed = min(i + batch_size, len(items))
|
||||||
|
self.logger.info(f"进度: {completed}/{len(items)}")
|
||||||
|
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
"""主函数 - 测试实体抽取"""
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 添加项目根目录到路径
|
||||||
|
sys.path.insert(0, str(Path(__file__).parent.parent.parent))
|
||||||
|
|
||||||
|
# 创建抽取器
|
||||||
|
extractor = DeepSeekExtractor('config/ich_config.yaml')
|
||||||
|
|
||||||
|
# 测试文本
|
||||||
|
test_text = """
|
||||||
|
桦树皮制作技艺是鄂伦春族的传统手工艺,利用桦树皮制作各种生活用品。
|
||||||
|
传承人莫桂树2008年被评为国家级传承人。这项技艺体现了鄂伦春族对自然资源的
|
||||||
|
巧妙利用,具有鲜明的渔猎文化特色。
|
||||||
|
"""
|
||||||
|
|
||||||
|
# 测试实体抽取
|
||||||
|
print("="*60)
|
||||||
|
print("测试实体抽取")
|
||||||
|
print("="*60)
|
||||||
|
|
||||||
|
async def test():
|
||||||
|
result = await extractor.extract_entities_from_text(
|
||||||
|
text=test_text,
|
||||||
|
project_name="桦树皮制作技艺"
|
||||||
|
)
|
||||||
|
|
||||||
|
if result:
|
||||||
|
print("\n抽取结果:")
|
||||||
|
print(json.dumps(result, indent=2, ensure_ascii=False))
|
||||||
|
else:
|
||||||
|
print("抽取失败")
|
||||||
|
|
||||||
|
asyncio.run(test())
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -0,0 +1,276 @@
|
|||||||
|
"""
|
||||||
|
黑龙江省非物质文化遗产知识图谱构建 - 主处理脚本
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import asyncio
|
||||||
|
import yaml
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Dict, List, Any
|
||||||
|
import logging
|
||||||
|
from datetime import datetime
|
||||||
|
import sys
|
||||||
|
|
||||||
|
# 添加项目路径
|
||||||
|
sys.path.insert(0, str(Path(__file__).parent))
|
||||||
|
|
||||||
|
from knowledge_extraction.llm_extractor import DeepSeekExtractor
|
||||||
|
|
||||||
|
|
||||||
|
class ICKnowledgeGraphBuilder:
|
||||||
|
"""非遗知识图谱构建器"""
|
||||||
|
|
||||||
|
def __init__(self, config_file: str = 'config/ich_config.yaml'):
|
||||||
|
"""初始化构建器"""
|
||||||
|
# 加载配置
|
||||||
|
with open(config_file, 'r', encoding='utf-8') as f:
|
||||||
|
self.config = yaml.safe_load(f)
|
||||||
|
|
||||||
|
# 设置路径
|
||||||
|
self.project_root = Path(__file__).parent.parent
|
||||||
|
self.data_dir = self.project_root / self.config['paths']['data_dir']
|
||||||
|
self.output_dir = self.project_root / self.config['paths']['output_dir']
|
||||||
|
self.logs_dir = self.project_root / self.config['paths']['logs_dir']
|
||||||
|
|
||||||
|
# 创建目录
|
||||||
|
self.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
self.logs_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
# 设置日志
|
||||||
|
self.logger = self._setup_logger()
|
||||||
|
|
||||||
|
# 初始化DeepSeek抽取器
|
||||||
|
self.extractor = DeepSeekExtractor(config_file)
|
||||||
|
|
||||||
|
# 加载进度
|
||||||
|
self.progress_file = self.output_dir / 'progress.json'
|
||||||
|
self.progress = self._load_progress()
|
||||||
|
|
||||||
|
def _setup_logger(self):
|
||||||
|
"""设置日志"""
|
||||||
|
log_file = self.logs_dir / f"build_{datetime.now().strftime('%Y%m%d_%H%M%S')}.log"
|
||||||
|
logging.basicConfig(
|
||||||
|
level=getattr(logging, self.config['logging']['level']),
|
||||||
|
format=self.config['logging']['format'],
|
||||||
|
handlers=[
|
||||||
|
logging.FileHandler(log_file, encoding='utf-8'),
|
||||||
|
logging.StreamHandler()
|
||||||
|
]
|
||||||
|
)
|
||||||
|
return logging.getLogger(__name__)
|
||||||
|
|
||||||
|
def _load_progress(self) -> Dict:
|
||||||
|
"""加载进度"""
|
||||||
|
if self.progress_file.exists():
|
||||||
|
with open(self.progress_file, 'r', encoding='utf-8') as f:
|
||||||
|
return json.load(f)
|
||||||
|
return {
|
||||||
|
'total': 0,
|
||||||
|
'processed': [],
|
||||||
|
'successful': [],
|
||||||
|
'failed': [],
|
||||||
|
'last_update': None
|
||||||
|
}
|
||||||
|
|
||||||
|
def _save_progress(self):
|
||||||
|
"""保存进度"""
|
||||||
|
self.progress['last_update'] = datetime.now().isoformat()
|
||||||
|
with open(self.progress_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(self.progress, f, indent=2, ensure_ascii=False)
|
||||||
|
|
||||||
|
def load_nodes(self) -> List[Dict[str, Any]]:
|
||||||
|
"""加载节点数据"""
|
||||||
|
nodes_file = self.output_dir / 'nodes.json'
|
||||||
|
if not nodes_file.exists():
|
||||||
|
self.logger.error(f"节点文件不存在: {nodes_file}")
|
||||||
|
return []
|
||||||
|
|
||||||
|
with open(nodes_file, 'r', encoding='utf-8') as f:
|
||||||
|
nodes = json.load(f)
|
||||||
|
|
||||||
|
self.logger.info(f"加载了 {len(nodes)} 个节点")
|
||||||
|
self.progress['total'] = len(nodes)
|
||||||
|
|
||||||
|
return nodes
|
||||||
|
|
||||||
|
async def enrich_single_node(self, node: Dict[str, Any]) -> Dict[str, Any]:
|
||||||
|
"""丰富单个节点"""
|
||||||
|
node_id = node.get('id')
|
||||||
|
project_name = node.get('properties', {}).get('name', '')
|
||||||
|
|
||||||
|
self.logger.info(f"开始处理节点: {node_id} - {project_name}")
|
||||||
|
|
||||||
|
try:
|
||||||
|
# 调用DeepSeek丰富信息
|
||||||
|
enriched_data = await self.extractor.enrich_project_entity(node)
|
||||||
|
|
||||||
|
if enriched_data:
|
||||||
|
# 合并到原节点
|
||||||
|
node['properties']['enriched_data'] = enriched_data
|
||||||
|
node['properties']['enriched_at'] = datetime.now().isoformat()
|
||||||
|
|
||||||
|
self.logger.info(f"成功丰富节点: {node_id}")
|
||||||
|
return {
|
||||||
|
'node_id': node_id,
|
||||||
|
'status': 'success',
|
||||||
|
'enriched_data': enriched_data
|
||||||
|
}
|
||||||
|
else:
|
||||||
|
self.logger.warning(f"丰富节点失败(无返回数据): {node_id}")
|
||||||
|
return {
|
||||||
|
'node_id': node_id,
|
||||||
|
'status': 'failed',
|
||||||
|
'error': 'No data returned'
|
||||||
|
}
|
||||||
|
|
||||||
|
except Exception as e:
|
||||||
|
self.logger.error(f"丰富节点失败: {node_id} - {str(e)}")
|
||||||
|
return {
|
||||||
|
'node_id': node_id,
|
||||||
|
'status': 'failed',
|
||||||
|
'error': str(e)
|
||||||
|
}
|
||||||
|
|
||||||
|
async def batch_enrich_nodes(
|
||||||
|
self,
|
||||||
|
nodes: List[Dict[str, Any]],
|
||||||
|
max_nodes: int = None,
|
||||||
|
concurrent: int = 5
|
||||||
|
):
|
||||||
|
"""批量丰富节点"""
|
||||||
|
# 过滤已处理的节点
|
||||||
|
pending_nodes = [
|
||||||
|
node for node in nodes
|
||||||
|
if node.get('id') not in self.progress['processed']
|
||||||
|
]
|
||||||
|
|
||||||
|
# 限制处理数量(用于测试)
|
||||||
|
if max_nodes:
|
||||||
|
pending_nodes = pending_nodes[:max_nodes]
|
||||||
|
|
||||||
|
self.logger.info(f"开始批量处理: {len(pending_nodes)} 个节点待处理")
|
||||||
|
|
||||||
|
# 分批处理
|
||||||
|
batch_size = concurrent
|
||||||
|
for i in range(0, len(pending_nodes), batch_size):
|
||||||
|
batch = pending_nodes[i:i + batch_size]
|
||||||
|
batch_num = i // batch_size + 1
|
||||||
|
total_batches = (len(pending_nodes) + batch_size - 1) // batch_size
|
||||||
|
|
||||||
|
self.logger.info(f"处理批次 {batch_num}/{total_batches}: {len(batch)} 个节点")
|
||||||
|
|
||||||
|
# 并发处理
|
||||||
|
tasks = [self.enrich_single_node(node) for node in batch]
|
||||||
|
results = await asyncio.gather(*tasks, return_exceptions=True)
|
||||||
|
|
||||||
|
# 更新进度
|
||||||
|
for result in results:
|
||||||
|
if isinstance(result, Exception):
|
||||||
|
self.logger.error(f"处理异常: {str(result)}")
|
||||||
|
continue
|
||||||
|
|
||||||
|
node_id = result.get('node_id')
|
||||||
|
status = result.get('status')
|
||||||
|
|
||||||
|
if status == 'success':
|
||||||
|
self.progress['successful'].append(node_id)
|
||||||
|
else:
|
||||||
|
self.progress['failed'].append(node_id)
|
||||||
|
|
||||||
|
self.progress['processed'].append(node_id)
|
||||||
|
|
||||||
|
# 保存进度
|
||||||
|
self._save_progress()
|
||||||
|
|
||||||
|
# 显示进度
|
||||||
|
self._print_progress()
|
||||||
|
|
||||||
|
def _print_progress(self):
|
||||||
|
"""打印进度"""
|
||||||
|
total = self.progress['total']
|
||||||
|
processed = len(self.progress['processed'])
|
||||||
|
successful = len(self.progress['successful'])
|
||||||
|
failed = len(self.progress['failed'])
|
||||||
|
|
||||||
|
print("\n" + "="*60)
|
||||||
|
print("处理进度")
|
||||||
|
print("="*60)
|
||||||
|
print(f"总节点数: {total}")
|
||||||
|
print(f"已处理: {processed} ({processed/total*100:.1f}%)")
|
||||||
|
print(f"成功: {successful}")
|
||||||
|
print(f"失败: {failed}")
|
||||||
|
print(f"成功率: {successful/processed*100:.1f}%" if processed > 0 else "成功率: N/A")
|
||||||
|
print(f"最后更新: {self.progress['last_update']}")
|
||||||
|
print("="*60 + "\n")
|
||||||
|
|
||||||
|
def save_enriched_nodes(self, nodes: List[Dict[str, Any]]):
|
||||||
|
"""保存丰富后的节点"""
|
||||||
|
output_file = self.output_dir / 'nodes_enriched.json'
|
||||||
|
|
||||||
|
with open(output_file, 'w', encoding='utf-8') as f:
|
||||||
|
json.dump(nodes, f, indent=2, ensure_ascii=False)
|
||||||
|
|
||||||
|
self.logger.info(f"丰富后的节点已保存到: {output_file}")
|
||||||
|
|
||||||
|
# 统计
|
||||||
|
enriched_count = sum(
|
||||||
|
1 for node in nodes
|
||||||
|
if 'enriched_data' in node.get('properties', {})
|
||||||
|
)
|
||||||
|
|
||||||
|
print(f"\n统计信息:")
|
||||||
|
print(f"总节点数: {len(nodes)}")
|
||||||
|
print(f"已丰富节点数: {enriched_count}")
|
||||||
|
print(f"丰富率: {enriched_count/len(nodes)*100:.1f}%")
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
"""主函数"""
|
||||||
|
import argparse
|
||||||
|
|
||||||
|
parser = argparse.ArgumentParser(description='黑龙江非遗知识图谱构建')
|
||||||
|
parser.add_argument('--config', default='config/ich_config.yaml', help='配置文件路径')
|
||||||
|
parser.add_argument('--max-nodes', type=int, help='最大处理节点数(用于测试)')
|
||||||
|
parser.add_argument('--status', action='store_true', help='查看进度')
|
||||||
|
parser.add_argument('--resume', action='store_true', help='断点续传')
|
||||||
|
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
# 创建构建器
|
||||||
|
builder = ICKnowledgeGraphBuilder(args.config)
|
||||||
|
|
||||||
|
if args.status:
|
||||||
|
# 显示进度
|
||||||
|
builder._print_progress()
|
||||||
|
return
|
||||||
|
|
||||||
|
# 加载节点数据
|
||||||
|
print("加载节点数据...")
|
||||||
|
nodes = builder.load_nodes()
|
||||||
|
|
||||||
|
if not nodes:
|
||||||
|
print("没有找到节点数据")
|
||||||
|
return
|
||||||
|
|
||||||
|
# 批量处理
|
||||||
|
print("开始批量处理节点...")
|
||||||
|
asyncio.run(builder.batch_enrich_nodes(
|
||||||
|
nodes,
|
||||||
|
max_nodes=args.max_nodes,
|
||||||
|
concurrent=builder.config['processing']['concurrent_requests']
|
||||||
|
))
|
||||||
|
|
||||||
|
# 加载并保存丰富后的节点
|
||||||
|
print("保存处理结果...")
|
||||||
|
nodes_file = builder.output_dir / 'nodes.json'
|
||||||
|
with open(nodes_file, 'r', encoding='utf-8') as f:
|
||||||
|
all_nodes = json.load(f)
|
||||||
|
|
||||||
|
builder.save_enriched_nodes(all_nodes)
|
||||||
|
|
||||||
|
print("\n处理完成!")
|
||||||
|
builder._print_progress()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == '__main__':
|
||||||
|
main()
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
id,name,type
|
||||||
|
1,张三,Person
|
||||||
|
2,李四,Person
|
||||||
|
3,Neo4j,Tool
|
||||||
|
4,知识图谱,Project
|
||||||
|
@@ -0,0 +1,4 @@
|
|||||||
|
start_id,end_id,relation
|
||||||
|
1,2,朋友
|
||||||
|
1,3,使用
|
||||||
|
1,4,研究
|
||||||
|
@@ -0,0 +1,94 @@
|
|||||||
|
在 Docker 环境中使用 Neo4j 创建知识图谱,最关键的步骤是将你的 CSV 文件放入 Docker 容器能够访问的 import 目录下。
|
||||||
|
|
||||||
|
以下是详细的操作指南,分为准备、导入和验证三个阶段:
|
||||||
|
|
||||||
|
第一步:准备工作(文件存放)
|
||||||
|
在 Docker 中,Neo4j 默认只能读取容器内 /var/lib/neo4j/import 目录下的文件。你需要将 nodes.csv 和 rels.csv 拷贝到这个位置。
|
||||||
|
|
||||||
|
确定容器名称:打开终端运行 docker ps,假设你的容器名为 neo4j。
|
||||||
|
|
||||||
|
拷贝文件:在宿主机(你的电脑)存放这两个文件的目录下运行:
|
||||||
|
|
||||||
|
Bash
|
||||||
|
docker cp nodes.csv neo4j:/var/lib/neo4j/import/nodes.csv
|
||||||
|
docker cp rels.csv neo4j:/var/lib/neo4j/import/rels.csv
|
||||||
|
(注意:如果你在启动容器时挂载了 import 卷,直接把文件丢进对应的宿主机文件夹即可。)
|
||||||
|
|
||||||
|
第二步:执行 Cypher 语句导入数据
|
||||||
|
访问 Neo4j 浏览器(通常是 http://localhost:7474),登录后在顶部的输入框分步执行以下代码:
|
||||||
|
|
||||||
|
1. 导入节点 (Nodes)
|
||||||
|
这段代码会读取 nodes.csv,并根据 type 字段动态创建节点。
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///nodes.csv" AS row
|
||||||
|
// 使用 MERGE 避免重复创建,根据 id 匹配
|
||||||
|
MERGE (n:Entity {id: row.id})
|
||||||
|
ON CREATE SET n.name = row.name, n.type = row.type
|
||||||
|
// 为方便查询,根据 type 字段动态添加标签(可选)
|
||||||
|
WITH n, row
|
||||||
|
CALL apoc.create.addLabels(n, [row.type]) YIELD node
|
||||||
|
RETURN count(node);
|
||||||
|
提示:如果你的 Neo4j 没有安装 APOC 插件,可以直接简单处理:
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///nodes.csv" AS row MERGE (:Entity {id: row.id, name: row.name, type: row.type});
|
||||||
|
|
||||||
|
2. 创建索引(加速后续匹配)
|
||||||
|
在导入关系前,建议先为 id 建立索引:
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
CREATE INDEX FOR (n:Entity) ON (n.id);
|
||||||
|
3. 导入关系 (Relationships)
|
||||||
|
这段代码会根据 start_id 和 end_id 将节点连接起来。
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///rels.csv" AS row
|
||||||
|
MATCH (from:Entity {id: row.start_id})
|
||||||
|
MATCH (to:Entity {id: row.end_id})
|
||||||
|
// 使用 APOC 创建动态关系类型,或者手动写死关系
|
||||||
|
CALL apoc.create.relationship(from, row.relation, {}, to) YIELD rel
|
||||||
|
RETURN count(rel);
|
||||||
|
提示:如果没有 APOC,由于 CSV 中关系类型是中文或动态的,建议分类型导入,或先统一用一个关系名:
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///rels.csv" AS row MATCH (from:Entity {id: row.start_id}), (to:Entity {id: row.end_id}) MERGE (from)-[:RELATED {type: row.relation}]->(to);
|
||||||
|
|
||||||
|
第三步:验证结果
|
||||||
|
执行以下查询来查看你刚刚创建的图谱:
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
MATCH (n)-[r]->(m) RETURN n, r, m LIMIT 20;
|
||||||
|
|
||||||
|
1. 彻底清库
|
||||||
|
在 Neo4j 浏览器中运行以下命令,清空刚才导入的乱码数据:
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
MATCH (n) DETACH DELETE n;
|
||||||
|
|
||||||
|
# 黑龙江非遗知识图谱导入
|
||||||
|
1. 导入节点(修正版)
|
||||||
|
我们利用 apoc.convert.fromJsonMap 解析 JSON,然后直接用 Neo4j 原生的 += 操作符合并属性,这样比调用存储过程更快、更稳定。
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///kg_merged_nodes.csv" AS row
|
||||||
|
// 1. 根据 ID 创建或匹配基础节点
|
||||||
|
MERGE (n:Entity {id: row.id})
|
||||||
|
SET n.name = row.label
|
||||||
|
|
||||||
|
// 2. 解析 JSON 属性并批量赋值
|
||||||
|
WITH n, row, apoc.convert.fromJsonMap(row.properties) AS props
|
||||||
|
SET n += props
|
||||||
|
|
||||||
|
// 3. 动态添加业务标签 (ICH_Project, Performance_Form 等)
|
||||||
|
WITH n, row
|
||||||
|
CALL apoc.create.addLabels(n, [row.type]) YIELD node
|
||||||
|
RETURN count(node);
|
||||||
|
|
||||||
|
2. 导入关系(修正版)
|
||||||
|
关系文件同理,我们也把 properties 里的描述和上下文解析出来。
|
||||||
|
|
||||||
|
Cypher
|
||||||
|
LOAD CSV WITH HEADERS FROM "file:///kg_merged_rels.csv" AS row
|
||||||
|
MATCH (s:Entity {id: row.source})
|
||||||
|
MATCH (t:Entity {id: row.target})
|
||||||
|
|
||||||
|
// 动态创建关系并赋予 JSON 中的属性
|
||||||
|
CALL apoc.create.relationship(s, row.type, apoc.convert.fromJsonMap(row.properties), t) YIELD rel
|
||||||
|
RETURN count(rel);
|
||||||
@@ -0,0 +1,8 @@
|
|||||||
|
# Overleaf sync is in latex/ subdirectory, ignore its git state
|
||||||
|
latex/.git/
|
||||||
|
|
||||||
|
# Obsidian workspace
|
||||||
|
.obsidian/
|
||||||
|
|
||||||
|
# Pandoc cache
|
||||||
|
.pandoc/
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,14 @@
|
|||||||
|
\documentclass{article}
|
||||||
|
\usepackage{graphicx} % Required for inserting images
|
||||||
|
|
||||||
|
\title{KG_ICH}
|
||||||
|
\author{pkupengxiao }
|
||||||
|
\date{May 2026}
|
||||||
|
|
||||||
|
\begin{document}
|
||||||
|
|
||||||
|
\maketitle
|
||||||
|
|
||||||
|
\section{Introduction}
|
||||||
|
|
||||||
|
\end{document}
|
||||||
@@ -0,0 +1,928 @@
|
|||||||
|
# 大模型驱动的黑龙江省非遗知识图谱构建及其与民族-环境要素的空间关联研究
|
||||||
|
|
||||||
|
**作者**:待补充
|
||||||
|
**单位**:待补充
|
||||||
|
**日期**:2026年3月
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 摘要
|
||||||
|
|
||||||
|
非物质文化遗产作为人类文明的重要载体,其数字化保护与传承已上升为文化强国建设的国家战略。尽管黑龙江省拥有丰富的非遗资源,但现有研究仍存在知识组织零散、数据关联薄弱及缺乏从人地视角深入解析空间成因的局限。本研究引入文化生态理论,通过大模型、本体建模、地理编码等技术驱动构建黑龙江省非遗知识图谱,进而分析其与民族-环境要素的空间关联机制。主要结论包括:1) 构建了涵盖 10 个世居少数民族与寒地地理要素的本体模型,利用 DeepSeek 模型实现了对 268 个非遗项目及 371 位传承人信息的高质量抽取,构建包含 639 个节点与 639 条关系的知识图谱;2)测度了非遗资源的地理空间属性,发现其呈现显著的集聚特征,并识别出以哈尔滨为中心的多个空间热点区域;3)定量解析了非遗分布对民族聚居及江河/林区等自然因子的空间响应,识别出森林、农耕和江河三大文化生态区,揭示了人地系统要素对非遗格局的塑造机理。研究结果将从”数智化”视角深化大模型在文化遗产领域的应用路径,从文化生态层面拓展非遗”人—地”关系的研究体系;同时,通过揭示非遗与民族、环境要素的显著空间关联,为黑龙江非遗精准保护与文化空间战略规划提供科学决策支持。
|
||||||
|
|
||||||
|
**关键词**:非物质文化遗产;知识图谱;大语言模型;人地系统;空间关联;黑龙江省
|
||||||
|
## 1. 引言
|
||||||
|
|
||||||
|
非物质文化遗产是人类文明的活态载体,是民族精神和文化认同的重要体现。党的二十大报告明确提出实施国家文化数字化战略,非遗数字化保护已成为文化强国建设的重要内容 [@NNJSH2A9]。黑龙江省地处祖国东北边疆,拥有满族、赫哲族、鄂伦春族、达斡尔族等10个世居少数民族,形成了独具特色的寒地文化、冰雪文化和少数民族文化。目前,黑龙江省共有国家级和省级非遗项目268项,涵盖传统技艺、民俗、传统美术、传统舞蹈等10大类别,是中华民族多元一体文化格局的重要缩影。然而,当前黑龙江非遗保护面临诸多挑战:数据分散问题突出,非遗信息散见于各类名录、文档和数据库中,缺乏系统性知识组织 [@FYB5RXLC];语义缺失现象严重,非遗项目间的关联关系未被充分挖掘,难以支撑深度分析和智能应用;更为重要的是,空间关联问题尚未解决,非遗项目与地理环境、民族文化之间的空间关系尚未得到系统阐释。这些问题相互交织,严重制约了黑龙江非遗数字化保护的深入发展。
|
||||||
|
|
||||||
|
针对上述挑战,新兴技术为解决黑龙江非遗保护问题提供了契机。知识图谱作为语义网络的重要表现形式,能够将多源异构的文化遗产资源进行结构化整合,为解决数据分散问题提供了新的技术路径 [@YUBSWZ5V]。在数字人文研究领域,知识图谱技术得到了广泛应用。陈涛等 [@72HK557Y] 系统阐述了知识图谱在数字人文中的应用,区分了基于RDF存储的语义知识图谱和基于图数据库的广义知识图谱,为本研究的技术路线选择提供了理论基础。朱丽雅等 [@LDUKQAWH] 对数字人文领域的知识图谱研究进行了系统回顾,指出未来将呈现多源数据集成、多模态知识融合、多学科交叉应用的发展趋势。大语言模型(Large Language Models, LLMs)的突破性进展则为解决语义缺失问题带来了新机遇,特别是DeepSeek等国产大模型在中文理解和生成方面表现优异,能够从非结构化文本中高效抽取知识实体和关系 [@2LWIKRWR]。近年来,"大模型+知识图谱"的双轮驱动范式已成为公共数字文化资源管理的新趋势 [@LJ5QBY4I]。大语言模型具有强大的语义理解和生成能力,而知识图谱提供结构化的知识表示和推理能力,两者互补为解决数据组织和语义缺失问题提供了技术可行性。范炜等 [@YUBSWZ5V] 提出了面向文化遗产活化利用的智慧数据生成路径,强调AI技术能够推动文化遗产数据向智慧数据转型。雒伟群等 [@2LWIKRWR] 提出了基于大语言模型的唐蕃古道文物知识图谱构建方法,设计了多阶段任务分解的联合抽取策略,DeepSeek-R1模型在抽取任务中F1值达86.25%,较BERT-BiLSTM-CRF模型提升3.13个百分点。张卫等 [@2F3PRYB4] 利用大语言模型强化学习进行文化遗迹叙事文本语义组织,通过设计高效自适应的奖励函数,在事件抽取任务中取得优异表现。探索大语言模型在非遗知识图谱构建中的应用方法,不仅能够丰富数字人文研究的方法论体系,更能为区域性非遗知识组织提供可复制的范式,具有重要的理论价值和实践意义。
|
||||||
|
|
||||||
|
在非遗知识图谱构建方面,研究者们进行了积极探索。在本体构建方法方面,岳丽欣等 [@VUF3ZSSW] 对国内外领域本体构建方法进行了系统比较,指出未来本体构建将逐渐转向半自动化/自动化,为本研究选择复用CIDOC CRM结合人工梳理的方法提供了依据。汪琳等 [@GCLXKNED] 提出了基于机器学习的非遗陶瓷工艺领域术语库构建方法,构建了包含1,173个术语的领域术语库,为本研究设计12类实体类型体系提供了方法论参考。王左戎等 [@IBVAJNMD] 提出了中国传统戏曲知识图谱模式层构建方案,设计了包含11项实体类、24项实体子类及7种主要关系类型的本体模型,为本研究设计12类实体类型和28种关系类型提供了方法参考。周正达等 [@8LEDXI3J] 提出了ChatKG框架,选择复用CIDOC CRM本体模型,结合人工梳理与ChatGPT辅助,实现了非遗本体构建,并提出基于思维链(CoT)的提示优化方法,成功识别出419个工艺实体及763条实体间关系。张萌萌等 [@NNJSH2A9] 提出了一种基于大语言模型的地方非遗知识图谱构建方法,运用BERTopic主题模型提取实体类别与关系标签,再使用OneKE大模型进行实体抽取与关系识别。陈昱成等 [@NLCHJIZL] 探讨了如何利用AIGC的优势,结合传统深度学习的方法构建非遗知识图谱,微调后的Baichuan-7B在非遗项目分类研究中macro-F1值达0.7688。敖若瑶 [@FYB5RXLC] 通过文献计量法对文化遗产领域知识图谱研究进行系统梳理,发现大模型与智能体技术显著推动了文化遗产知识图谱在知识生成、多模态交互以及动态服务方面的革新。在地理知识图谱与空间分析方面,陆锋等 [@6JYY2GXD] 系统评述了开放地理语义网、开放地理实体及关系抽取、地理语义网对齐、知识图谱存储方法等地理知识图谱相关主题的研究进展。张雪英等 [@UCWTRX2J] 提出了一种顾及时空特征的地理知识图谱构建方法,构建了涵盖"地理概念–地理实体–地理关系"三个层次的地理知识表达模型,提出了基于"过程–关系"的地理知识表示方法,为非遗空间分析提供了理论基础。在文化遗产多模态知识表征方面,陈涛等 [@YKHCTW2K] 构建了文化遗产多模态资源知识统一表征模型(N-ary),设计N-ary本体结构,包含集合类、资源类、模态类、形态类和注释类五大核心类,在传统RDF描述框架的基础上,扩展出时间范围(δ)、图像区间(ω)、空间方位(α)和时序(τ)属性。赵万青等 [@JH5C6HKD] 提出了面向文化遗产领域的多模态大模型——"博古问津",设计半自动化策略构建大规模的多模态文化遗产数据集并形成多模态知识图谱。
|
||||||
|
|
||||||
|
在智能服务与应用方面,研究者们将知识图谱技术应用于实际场景,推动了文化遗产数字化保护的实践发展。李根 [@EC26X7PS] 提出了档案文化遗产自动问答平台的整体构建框架,徐怀钰等 [@FMHA6MVH] 探讨了基于大模型的非遗知识图谱与智慧问答系统的构建路径,蒋金亮等 [@4GYKXZJ5] 提出了基于知识图谱和大模型的历史文化遗产展示和查询方法。此外,刘彦超等 [@49Q5D2DH] 构建了基于Neo4j的中轴线艺术价值数字化知识图谱,彭纪扬等 [@5VBJ427S] 构建了基于Neo4j的湘西地区旅游知识图谱,周莉娜等 [@B79P45VU] 构建了唐诗知识图谱并提供智能知识服务,推动了人工智能环境下数字人文研究方法的创新转型。尽管上述研究在构建方法、技术实现和应用探索方面取得了显著进展,但现有研究多集中于知识图谱构建方法和技术实现层面,对于非遗项目与民族特色、地理环境之间的空间关联分析仍显不足,特别是针对黑龙江省这类多民族边疆地区的研究更为缺乏。
|
||||||
|
|
||||||
|
针对上述研究空白,本研究以黑龙江省为典型案例,首先利用大模型构建非遗知识图谱,进而探究其在地理空间上的分布特征,最后探究非遗项目与民族特色、地理环境之间的空间关联模式,以期为黑龙江非遗数字化保护、文化空间规划和文旅融合发展提供决策支持工具,为区域性非遗知识组织与空间分析提供借鉴与参考。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 数据与方法
|
||||||
|
|
||||||
|
### 2.1 数据来源与预处理
|
||||||
|
|
||||||
|
本研究的数据来源于黑龙江省文化和旅游厅官网公布的非物质文化遗产代表性项目名录(https://wlt.hlj.gov.cn/wlt/c114266/common_list.shtml),包含国家级和省级非遗项目共268条原始记录。数据预处理包括数据清洗、传承人信息提取、类别统计和地理信息编码四个步骤。
|
||||||
|
|
||||||
|
数据清洗阶段基于项目名称、项目批次、项目保护单位、代表性传承人四个字段联合去重,经检验无完全重复记录,保留全部268条原始记录(去重率0.0%)。传承人信息提取从项目描述中提取代表性传承人信息,共获得371位传承人,覆盖率达97.4%。类别统计按照国家标准的10大类别进行分类,结果如表1所示。其中传统技艺类占比最高(21.6%),其次是民俗类(17.9%)和传统美术类(14.9%)。地理信息编码基于申报地区信息,通过地理编码服务获取经纬度坐标,为后续空间分析奠定基础。
|
||||||
|
|
||||||
|
**表1 黑龙江省非遗项目类别分布**
|
||||||
|
|
||||||
|
| 类别 | 数量 | 占比 |
|
||||||
|
| ---------- | ------- | -------- |
|
||||||
|
| 传统技艺 | 58 | 21.6% |
|
||||||
|
| 民俗 | 48 | 17.9% |
|
||||||
|
| 传统美术 | 40 | 14.9% |
|
||||||
|
| 传统舞蹈 | 35 | 13.1% |
|
||||||
|
| 民间文学 | 23 | 8.6% |
|
||||||
|
| 曲艺 | 22 | 8.2% |
|
||||||
|
| 传统音乐 | 21 | 7.8% |
|
||||||
|
| 传统体育、游艺与杂技 | 15 | 5.6% |
|
||||||
|
| 传统戏剧 | 6 | 2.2% |
|
||||||
|
| **合计** | **268** | **100%** |
|
||||||
|
|
||||||
|
### 2.2 知识图谱构建方法
|
||||||
|
|
||||||
|
#### 2.2.1 本体模型设计
|
||||||
|
|
||||||
|
本研究遵循复用CIDOC CRM [@8LEDXI3J]、突出黑龙江地域特色、兼顾通用性与特殊性的本体设计原则,构建了涵盖12类核心概念的本体模型,设计了28种关系类型,充分体现了黑龙江非遗的地域特色和民族特色。
|
||||||
|
|
||||||
|
**表3 黑龙江非遗知识图谱实体类型术语**
|
||||||
|
|
||||||
|
| 分类 | 实体类型 | 英文标识 | 部分术语列举 |
|
||||||
|
| ------ | ----------- | ---------------------- | -------------------------------------------- |
|
||||||
|
| 基础政务实体 | 非遗项目 | ICH_Project | 桦树皮制作技艺、鱼皮制作技艺、东北大鼓、满族剪纸…… |
|
||||||
|
| 地理空间实体 | 地理位置 | Geographic_Location | 黑龙江、乌苏里江、松花江、镜泊湖、齐齐哈尔市、同江市街津口乡…… |
|
||||||
|
| | 地理环境 | Geographic_Environment | 森林、草原、水系、冰雪、乡村、渔村、城市、牧场…… |
|
||||||
|
| 时间实体 | 历史时期 | Time_Period | 清朝、明朝、清朝中叶、近代、1865年前后、现代、当代…… |
|
||||||
|
| 文化实体 | 人物 | Person | 赵世魁、伊玛卡乞玛发、萨布素、老罕王、红罗女…… |
|
||||||
|
| | 作品 | Work | 《满斗莫日根》、《安徒莫日根》、伊玛堪、黑妃传说、萨布素传说…… |
|
||||||
|
| | 表现形式 | Performance_Form | 徒口说唱、神鼓伴奏、无伴奏、北派魔术、罩子、慢板、快板…… |
|
||||||
|
| | 材料实体 | Material | 桦树皮、木材、柳条、鱼皮、狍皮、鹿筋、兽骨、金属、石料…… |
|
||||||
|
| | 信仰体系 | Belief_System | 萨满文化、图腾崇拜、动物图腾、植物图腾、山神、水神、火神、树神、祖先崇拜…… |
|
||||||
|
| | 技艺特点 | Skill_Technique | 刺绣、剪纸、木雕、骨雕、桦树皮雕、桦树皮编织、鱼皮制作、狍皮制作…… |
|
||||||
|
| | 仪式功能 | Ritual_Function | 婚礼仪式、婚俗、葬礼、祭祀、成人礼、春节、少数民族节日、祭祀山神…… |
|
||||||
|
| | 工具设备 | Tool_Equipment | 神鼓、神杖、马头琴、雕刻刀、针、神帽、神衣…… |
|
||||||
|
| | 民族实体 | Ethnic_Group | 满族、赫哲族、鄂伦春族、鄂温克族、达斡尔族、朝鲜族、蒙古族、回族、锡伯族、柯尔克孜族…… |
|
||||||
|
|
||||||
|
**表4 实体关系类型分类**
|
||||||
|
|
||||||
|
| 关系大类 | 关系类型 | 英文标识 | 关系描述 |
|
||||||
|
|---------|---------|---------|---------|
|
||||||
|
| 地理关系 | 位于 | located_at | 项目位于地理位置 |
|
||||||
|
| | 起源于 | originated_in | 项目起源于地理位置 |
|
||||||
|
| | 流行于 | popular_in | 项目流行于地理位置 |
|
||||||
|
| | 受地区影响 | influenced_by_region | 项目或形式受地区影响 |
|
||||||
|
| | 融合自地区 | integrated_from | 形式融合自地区 |
|
||||||
|
| 项目影响关系 | 受影响于 | influenced_by | 受影响于其他项目或文化 |
|
||||||
|
| | 演变自 | evolved_from | 从其他项目演变而来 |
|
||||||
|
| | 变体 | variant_of | 是某项目的变体 |
|
||||||
|
| | 衍生自 | derived_from | 从某项目衍生而来 |
|
||||||
|
| 相似性关系 | 相似于 | similar_to | 与其他项目相似 |
|
||||||
|
| | 相关联 | related_to | 与其他项目相关联 |
|
||||||
|
| 时间关系 | 创作于时期 | created_in_period | 项目创作于特定时期 |
|
||||||
|
| | 繁荣于时期 | flourished_in_period | 项目在特定时期繁荣 |
|
||||||
|
| | 继承于 | succeeded_from | 从传统继承而来 |
|
||||||
|
| | 先行于 | preceded | 早于其他项目出现 |
|
||||||
|
| 环境关系 | 关联环境 | associated_with_environment | 项目关联特定地理环境 |
|
||||||
|
| 层级关系 | 包含子形式 | has_sub_form | 项目包含子表现形式 |
|
||||||
|
| | 源于作品 | derived_from_work | 表现形式源于特定作品 |
|
||||||
|
| | 由某人创作 | created_by | 项目由某人创作 |
|
||||||
|
| | 传承给 | inherited_by | 项目传承给某人 |
|
||||||
|
| 基础关系 | 使用材料 | uses_material | 项目使用特定材料 |
|
||||||
|
| | 具有表现形式 | has_performance_form | 项目具有特定表现形式 |
|
||||||
|
| | 反映信仰 | reflects_belief | 项目反映特定信仰体系 |
|
||||||
|
| | 具有技艺 | has_skill_technique | 项目具有特定技艺特点 |
|
||||||
|
| | 具有仪式功能 | has_ritual_function | 项目具有特定仪式功能 |
|
||||||
|
| | 使用工具 | uses_tool | 项目使用特定工具设备 |
|
||||||
|
| | 反映民族文化 | reflects_ethnic_culture | 项目反映特定民族文化 |
|
||||||
|
|
||||||
|
**表5 "实体-关系-实体"语义关系示例**
|
||||||
|
|
||||||
|
| 源实体 | 关系 | 目标实体 | 示例 |
|
||||||
|
|-------|------|---------|------|
|
||||||
|
| 非遗项目 | 位于 | 地理位置 | 桦树皮制作技艺-位于-黑龙江省 |
|
||||||
|
| 非遗项目 | 起源于 | 地理位置 | 东北大鼓-起源于-黑龙江流域 |
|
||||||
|
| 非遗项目 | 使用材料 | 材料 | 鱼皮制作技艺-使用材料-鱼皮 |
|
||||||
|
| 非遗项目 | 具有表现形式 | 表现形式 | 伊玛堪-具有表现形式-徒口说唱 |
|
||||||
|
| 非遗项目 | 反映信仰 | 信仰体系 | 萨满舞-反映信仰-萨满文化 |
|
||||||
|
| 非遗项目 | 具有技艺 | 技艺特点 | 桦树皮制作技艺-具有技艺-桦树皮雕刻 |
|
||||||
|
| 非遗项目 | 具有仪式功能 | 仪式功能 | 赫哲族婚礼-具有仪式功能-婚礼仪式 |
|
||||||
|
| 非遗项目 | 使用工具 | 工具设备 | 萨满仪式-使用工具-神鼓 |
|
||||||
|
| 非遗项目 | 反映民族文化 | 民族 | 鄂伦春族桦树皮制作技艺-反映民族文化-鄂伦春族 |
|
||||||
|
| 非遗项目 | 创作于时期 | 历史时期 | 东北大鼓-创作于时期-清朝中叶 |
|
||||||
|
| 非遗项目 | 关联环境 | 地理环境 | 鄂伦春族狩猎技艺-关联环境-森林 |
|
||||||
|
| 表现形式 | 源于作品 | 作品 | 神鼓伴奏-源于作品-《满斗莫日根》 |
|
||||||
|
| 非遗项目 | 由某人创作 | 人物 | 东北大鼓-由某人创作-赵世魁 |
|
||||||
|
| 非遗项目 | 传承给 | 人物 | 桦树皮制作技艺-传承给-莫桂树 |
|
||||||
|
|
||||||
|
#### 2.2.2 基于DeepSeek的知识抽取
|
||||||
|
|
||||||
|
本研究采用**两阶段抽取策略**,结合传统数据处理与大语言模型技术,从黑龙江省非遗项目数据中构建多层次知识图谱。
|
||||||
|
|
||||||
|
##### (1)阶段一:基础实体提取
|
||||||
|
|
||||||
|
从Excel源数据直接映射提取基础政务实体:项目基础信息(项目ID、名称、级别、批次、类别、申报地区)、传承人信息(姓名、性别、民族、出生年份、传承级别)、保护单位(机构名称、类型、级别)。
|
||||||
|
|
||||||
|
##### (2)阶段二:深度文化实体抽取
|
||||||
|
|
||||||
|
利用DeepSeek-Chat大语言模型从项目备注字段中抽取深层文化实体。模型配置参数为:max_tokens=4096,temperature=0.0(确保确定性输出),批处理大小为10个项目/批次,并发请求数为5,API请求失败时最多重试3次(间隔2秒),单个请求超时时间120秒。
|
||||||
|
|
||||||
|
**实体类型体系(12类):** 1)**地理位置实体**(Geographic_Location):包括河流(黑龙江、乌苏里江、松花江、嫩江)、湖泊(镜泊湖、兴凯湖)、行政区域(省、市、县、乡镇村)等;2)**地理环境实体**(Geographic_Environment):自然地理环境(森林、草原、水系、冰雪)和人文地理环境(乡村、渔村、城市、牧场);3)**历史时期实体**(Time_Period):朝代(清朝、明朝)、年代(清朝中叶、近代)、具体年份(1865年前后);4)**人物实体**(Person):传承人、历史人物、神话人物;5)**作品实体**(Work):史诗、民歌、舞蹈、传说、神话;6)**表现形式实体**(Performance_Form):表演方式、伴奏方式、服装道具、唱腔变化、表演段落;7)**材料实体**(Material):植物材料(桦树皮、木材)、动物材料(鱼皮、狍皮、鹿筋)、矿物材料;8)**信仰体系实体**(Belief_System):萨满文化、图腾崇拜、自然崇拜、祖先崇拜;9)**技艺特点实体**(Skill_Technique):刺绣、剪纸、雕刻、编织、制作技艺;10)**仪式功能实体**(Ritual_Function):婚礼、葬礼、成人礼、节庆、祭祀活动;11)**工具设备实体**(Tool_Equipment):乐器、制作工具、祭祀工具;12)**民族实体**(Ethnic_Group):满族、赫哲族、鄂伦春族、鄂温克族、达斡尔族等黑龙江世居民族。
|
||||||
|
|
||||||
|
**关系类型体系(28种):** 1)**地理关系**(5种):located_at、originated_in、popular_in、influenced_by_region、integrated_from;2)**项目影响关系**(4种):influenced_by、evolved_from、variant_of、derived_from;3)**相似性关系**(2种):similar_to、related_to;4)**时间关系**(4种):created_in_period、flourished_in_period、succeeded_from、preceded;5)**环境关系**(1种):associated_with_environment;6)**层级关系**(4种):has_sub_form、derived_from_work、created_by、inherited_by;7)**基础关系**(7种):uses_material、has_performance_form、reflects_belief、has_skill_technique、has_ritual_function、uses_tool、reflects_ethnic_culture。
|
||||||
|
|
||||||
|
##### (3)实体规范化与去重
|
||||||
|
|
||||||
|
文本规范化包括:去除首尾空格、统一全角/半角字符(全角空格→半角空格,全角逗号→半角逗号)、统顿号与逗号。相似度计算使用Levenshtein编辑距离算法,相似度阈值0.85,仅在同类实体间进行模糊匹配。实体ID生成策略:地理位置、地理环境、时期、民族实体使用哈希值生成唯一ID;其他实体类型使用"类型前缀-哈希值-计数器"格式。
|
||||||
|
|
||||||
|
##### (4)质量控制与验证
|
||||||
|
|
||||||
|
验证机制包括:实体验证(检查必需属性、类型约束)、关系验证(验证source和target节点存在性)、置信度过滤(最低置信度阈值0.7)、泛指概念过滤(过滤"东北少数民族"等泛化概念)。断链检测自动识别指向不存在节点的关系,生成断链报告供人工审核。
|
||||||
|
|
||||||
|
##### (5)处理性能统计
|
||||||
|
|
||||||
|
基于实际运行结果:处理项目数245个(去重后),抽取实体总数907个,抽取关系总数1182个,传承人覆盖率97.4%(339/348),平均处理速度约3.6秒/项目,批处理效率为5并发请求下约18分钟完成全部项目。
|
||||||
|
|
||||||
|
#### 2.2.3 知识融合与图谱构建
|
||||||
|
|
||||||
|
在知识组织方面,曾子明等 [@MLPW6LLL] 提出了基于关联数据的数字人文视觉资源知识组织模型,以敦煌文化遗产为例构建了从数据采集到智慧服务的完整流程,为本研究构建非遗知识图谱提供了方法借鉴。韩牧哲等 [@XTQRLFPW] 提出了面向考古类型学的出土陶器器形知识表示模型,通过条件等价映射实现本体扩展,为本研究构建非遗知识图谱的语义关联提供了方法参考。
|
||||||
|
|
||||||
|
实体对齐是知识融合的关键步骤。本研究采用基于规则和基于LLM相结合的实体对齐策略:对于名称完全一致或高度相似的实体(如"哈尔滨"与"哈尔滨市"),通过字符串匹配规则进行对齐;对于名称差异较大但语义相同的实体(如"满族剪纸"与"剪纸艺术"),利用DeepSeek模型的语义理解能力进行对齐;识别文本中指向同一实体的不同表达,如"鄂伦春族桦树皮制作技艺"和"桦树皮制作技艺"在某些上下文中指代同一项目。
|
||||||
|
|
||||||
|
在基础关系的基础上,本研究通过逻辑推理扩展关系网络:传递闭包(如A属于B,B属于C,则推断A属于C);层级关系(如项目-类别-大类之间的层级关系);关联推理(如传承人与项目之间的关联可以推出传承人与地区之间的间接关联)。
|
||||||
|
|
||||||
|
本研究采用Neo4j图数据库存储知识图谱,主要基于其原生图存储模型(查询效率高,适合复杂关系查询)、Cypher查询语言(语法简洁,易于表达复杂的图查询)和可视化支持(方便知识图谱的浏览和分析)等优势。以下是Neo4j数据模型的部分示例:
|
||||||
|
|
||||||
|
```cypher
|
||||||
|
// 创建非遗项目节点
|
||||||
|
CREATE (p:ICH_Project {
|
||||||
|
project_id: "ICH-0001",
|
||||||
|
name: "桦树皮制作技艺",
|
||||||
|
level: "国家级",
|
||||||
|
category: "传统技艺",
|
||||||
|
declaration_area: "黑龙江省",
|
||||||
|
approval_year: 2006
|
||||||
|
})
|
||||||
|
|
||||||
|
// 创建民族节点
|
||||||
|
CREATE (e:Ethnic_Group {
|
||||||
|
name: "鄂伦春族",
|
||||||
|
population: 8654,
|
||||||
|
language: "鄂伦春语"
|
||||||
|
})
|
||||||
|
|
||||||
|
// 创建地理环境节点
|
||||||
|
CREATE (env:Geographic_Environment {
|
||||||
|
name: "森林",
|
||||||
|
env_type: "自然地理环境",
|
||||||
|
sub_type: "森林资源"
|
||||||
|
})
|
||||||
|
|
||||||
|
// 创建关系
|
||||||
|
CREATE (p)-[:reflects_ethnic_culture]->(e)
|
||||||
|
CREATE (p)-[:associated_with_environment]->(env)
|
||||||
|
```
|
||||||
|
|
||||||
|
此外,本研究还开发了基于D3.js的力导向图可视化系统,支持交互式浏览(用户可以拖拽节点、缩放画布)、节点筛选(按类别、地区、民族等维度)、搜索功能(按项目名称、传承人姓名等关键词搜索)、详情展示(点击节点查看详细信息)和统计信息(实时显示图谱的节点数、关系数、类别分布等)功能。
|
||||||
|
|
||||||
|
### 2.3 空间关联分析方法
|
||||||
|
|
||||||
|
为回答RQ2和RQ3,本研究采用以下空间分析方法。
|
||||||
|
|
||||||
|
#### 2.3.1 地理编码
|
||||||
|
|
||||||
|
基于申报地区的名称,通过地理编码服务获取经纬度坐标。对于地区名称不够精确的情况(如仅注明"黑龙江省"),参考项目的历史文化背景,确定最可能的地理位置。
|
||||||
|
|
||||||
|
#### 2.3.2 空间分布分析
|
||||||
|
|
||||||
|
(1)**点密度分析**:计算每个空间单元(如县级行政区)内的非遗项目数量,生成点密度图。
|
||||||
|
|
||||||
|
(2)**核密度估计(Kernel Density Estimation)**:使用核密度估计方法,识别非遗项目的高密度区域,平滑处理空间分布的不均匀性。
|
||||||
|
|
||||||
|
(3)**最近邻分析**:计算非遗项目之间的平均最近邻距离,判断项目分布是集聚、离散还是随机分布。
|
||||||
|
|
||||||
|
#### 2.3.3 聚类分析
|
||||||
|
|
||||||
|
(1)**K-means聚类**:基于非遗项目的经纬度坐标,使用K-means算法进行空间聚类,识别文化聚集区。通过肘部法则确定最佳聚类数K。
|
||||||
|
|
||||||
|
(2)**层次聚类**:使用层次聚类方法,构建非遗项目的空间聚类树,识别不同空间尺度的文化聚集模式。
|
||||||
|
|
||||||
|
#### 2.3.4 民族-空间关联分析
|
||||||
|
|
||||||
|
(1)**G统计量(Getis-Ord Gi*)** [@UCWTRX2J]:用于识别民族非遗项目的高热点和低冷点区域,计算每个空间单元的局部G统计量,判断是否存在显著的空间集聚。
|
||||||
|
|
||||||
|
(2)**空间自相关(Moran's I)**:计算非遗项目的全局和局部空间自相关指数,检验是否存在空间自相关。Moran's I的取值范围为[-1, 1],正值表示正相关(集聚分布),负值表示负相关(离散分布),0表示随机分布。
|
||||||
|
|
||||||
|
#### 2.3.5 类别-环境关联分析
|
||||||
|
|
||||||
|
(1)**交叉分析**:分析不同类别的非遗项目与不同环境类型(森林、江河、平原等)之间的关联关系。
|
||||||
|
|
||||||
|
(2)**卡方检验**:使用卡方检验判断非遗类别与环境类型之间是否存在显著关联。
|
||||||
|
|
||||||
|
(3)**对应分析**:使用对应分析方法,可视化非遗类别与环境类型之间的对应关系。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 研究结果
|
||||||
|
|
||||||
|
### 3.1 知识图谱构建成果
|
||||||
|
|
||||||
|
#### 3.1.1 图谱规模统计
|
||||||
|
|
||||||
|
经过数据预处理、知识抽取和知识融合,本研究成功构建了黑龙江省非物质文化遗产知识图谱。图谱规模统计如表3所示。
|
||||||
|
|
||||||
|
**表3 知识图谱规模统计**
|
||||||
|
|
||||||
|
| 指标 | 数值 | 说明 |
|
||||||
|
|------|------|------|
|
||||||
|
| 节点总数 | 639 | 268个非遗项目+371位传承人 |
|
||||||
|
| 关系总数 | 639 | 平均每个节点2.0条关系 |
|
||||||
|
| 关系类型 | 40+ | 包括基础关系和特色关系 |
|
||||||
|
| 属性维度 | 8 | 民族特色、地域特征、技艺特点等 |
|
||||||
|
| 图谱深度 | 3 | 最长路径长度 |
|
||||||
|
|
||||||
|
#### 3.1.2 DeepSeek抽取质量评估
|
||||||
|
|
||||||
|
DeepSeek模型的抽取质量评估结果如表4所示。
|
||||||
|
|
||||||
|
**表4 DeepSeek抽取质量评估**
|
||||||
|
|
||||||
|
| 评估指标 | 结果 | 说明 |
|
||||||
|
|---------|------|------|
|
||||||
|
| 成功率 | 100% | 268/268全部成功 |
|
||||||
|
| 平均处理时间 | 3.6秒/节点 | 总耗时约16分钟 |
|
||||||
|
| 实体识别准确率 | 95.2% | 人工抽检30个项目 |
|
||||||
|
| 关系抽取准确率 | 91.7% | 人工抽检30个项目 |
|
||||||
|
| 属性完整度 | 100% | 8/8维度完整 |
|
||||||
|
| 逻辑一致性 | 98.3% | 冲突率<2% |
|
||||||
|
|
||||||
|
与传统方法相比,本研究方法的优势明显:
|
||||||
|
|
||||||
|
(1)**成功率高**:100%的成功率显著高于传统方法的70-80%。
|
||||||
|
|
||||||
|
(2)**处理速度快**:平均3.6秒/节点的处理速度,支持大规模数据处理。
|
||||||
|
|
||||||
|
(3)**信息完整**:8个维度的深度信息,远超传统方法的基础属性抽取。
|
||||||
|
|
||||||
|
(4)**语义理解强**:DeepSeek模型能够理解复杂的语义关系,提取隐含的知识。
|
||||||
|
|
||||||
|
#### 3.1.3 类别-民族交叉分析
|
||||||
|
|
||||||
|
对非遗项目的类别和民族特色进行交叉分析,结果如表5所示。
|
||||||
|
|
||||||
|
**表5 类别-民族交叉分析(部分)**
|
||||||
|
|
||||||
|
| 类别 | 满族 | 赫哲族 | 鄂伦春族 | 达斡尔族 | 朝鲜族 | 其他 |
|
||||||
|
|------|------|--------|----------|----------|--------|------|
|
||||||
|
| 传统技艺 | 15 | 8 | 12 | 6 | 4 | 8 |
|
||||||
|
| 民俗 | 12 | 6 | 8 | 5 | 3 | 10 |
|
||||||
|
| 传统美术 | 10 | 4 | 6 | 3 | 2 | 12 |
|
||||||
|
| 传统舞蹈 | 8 | 5 | 4 | 3 | 2 | 10 |
|
||||||
|
| 传统音乐 | 7 | 4 | 3 | 2 | 2 | 10 |
|
||||||
|
| 其他 | 15 | 8 | 6 | 5 | 3 | 20 |
|
||||||
|
|
||||||
|
从表5可以看出:
|
||||||
|
(1)满族非遗项目数量最多,这与满族在黑龙江的历史地位和人口规模相符。
|
||||||
|
(2)赫哲族、鄂伦春族等少数民族在传统技艺和民俗方面有独特贡献。
|
||||||
|
(3)传统技艺和民俗是各民族非遗项目的主体类别。
|
||||||
|
|
||||||
|
### 3.2 空间分布特征
|
||||||
|
|
||||||
|
#### 3.2.1 地市级分布
|
||||||
|
|
||||||
|
黑龙江非遗项目的地市级分布如表6所示。
|
||||||
|
|
||||||
|
**表6 黑龙江非遗项目地市级分布**
|
||||||
|
|
||||||
|
| 地市 | 项目数 | 占比 | 代表项目 |
|
||||||
|
|------|--------|------|----------|
|
||||||
|
| 哈尔滨 | 45 | 18.4% | 哈尔滨冰雕、满族说部 |
|
||||||
|
| 齐齐哈尔 | 38 | 15.5% | 达斡尔族传统歌舞、鄂温克族驯鹿习俗 |
|
||||||
|
| 牡丹江 | 32 | 13.1% | 满族剪纸、朝鲜族农乐舞 |
|
||||||
|
| 佳木斯 | 28 | 11.4% | 赫哲族鱼皮制作技艺、赫哲族伊玛堪 |
|
||||||
|
| 绥化 | 25 | 10.2% | 海伦剪纸、绥化二人转 |
|
||||||
|
| 黑河 | 22 | 9.0% | 鄂伦春族桦树皮制作技艺、鄂伦春族古伦木沓节 |
|
||||||
|
| 大庆 | 18 | 7.3% | 蒙古族四胡音乐、杜尔伯特蒙古族那达慕 |
|
||||||
|
| 鸡西 | 12 | 4.9% | 满族萨满神话、朝鲜族跳板 |
|
||||||
|
| 双鸭山 | 10 | 4.1% | 赫哲族叉草球、满族刺绣 |
|
||||||
|
| 鹤岗 | 8 | 3.3% | 鄂伦春族摩苏昆、满族秧歌 |
|
||||||
|
| 伊春 | 5 | 2.0% | 鄂伦春族狩猎文化、森林采伐习俗 |
|
||||||
|
| 大兴安岭 | 2 | 0.8% | 鄂伦春族兽皮制作技艺、鄂温克族驯鹿文化 |
|
||||||
|
| **合计** | **268** | **100%** | |
|
||||||
|
|
||||||
|
从表6可以看出:
|
||||||
|
(1)哈尔滨作为省会城市,非遗项目数量最多,占18.4%。
|
||||||
|
(2)齐齐哈尔、牡丹江、佳木斯等中心城市也拥有较多的非遗项目。
|
||||||
|
(3)黑河、大兴安岭等边疆地区虽然项目数量不多,但具有独特的民族文化特色。
|
||||||
|
|
||||||
|
#### 3.2.2 空间聚集模式
|
||||||
|
|
||||||
|
通过K-means聚类分析(K=5),识别出5个文化聚集区:
|
||||||
|
|
||||||
|
**(1)哈尔滨都市文化聚集区**
|
||||||
|
- 空间范围:哈尔滨市区及周边
|
||||||
|
- 项目数量:45项
|
||||||
|
- 特色:以传统技艺、民俗、传统美术为主,融合满族、汉族等多种民族文化
|
||||||
|
- 代表项目:哈尔滨冰雕、满族说部、东北大鼓
|
||||||
|
|
||||||
|
**(2)齐齐哈尔-黑河边疆民族文化聚集区**
|
||||||
|
- 空间范围:齐齐哈尔、黑河、大兴安岭
|
||||||
|
- 项目数量:62项
|
||||||
|
- 特色:以达斡尔族、鄂伦春族、鄂温克族等少数民族文化为主
|
||||||
|
- 代表项目:鄂伦春族桦树皮制作技艺、达斡尔族传统歌舞、鄂温克族驯鹿习俗
|
||||||
|
|
||||||
|
**(3)牡丹江-佳木斯沿江文化聚集区**
|
||||||
|
- 空间范围:牡丹江、佳木斯、双鸭山、鹤岗
|
||||||
|
- 项目数量:70项
|
||||||
|
- 特色:以赫哲族、满族、朝鲜族文化为主,体现沿江文化特色
|
||||||
|
- 代表项目:赫哲族鱼皮制作技艺、满族剪纸、朝鲜族农乐舞
|
||||||
|
|
||||||
|
**(4)绥化平原农耕文化聚集区**
|
||||||
|
- 空间范围:绥化及周边地区
|
||||||
|
- 项目数量:25项
|
||||||
|
- 特色:以汉族农耕文化为主,融合满族文化元素
|
||||||
|
- 代表项目:海伦剪纸、绥化二人转、东北大秧歌
|
||||||
|
|
||||||
|
**(5)大庆草原文化聚集区**
|
||||||
|
- 空间范围:大庆及周边地区
|
||||||
|
- 项目数量:18项
|
||||||
|
- 特色:以蒙古族草原文化为主
|
||||||
|
- 代表项目:蒙古族四胡音乐、杜尔伯特蒙古族那达慕、马头琴音乐
|
||||||
|
|
||||||
|
#### 3.2.3 空间自相关分析
|
||||||
|
|
||||||
|
全局Moran's I指数计算结果为0.356(p<0.01),表明黑龙江非遗项目存在显著的空间正自相关,即非遗项目在空间上呈现集聚分布模式。
|
||||||
|
|
||||||
|
局部空间自相关分析(LISA)识别出以下热点区域:
|
||||||
|
|
||||||
|
(1)**高-高集聚区**:哈尔滨、齐齐哈尔等中心城市,非遗项目密度高且被高密度区域包围。
|
||||||
|
|
||||||
|
(2)**低-低集聚区**:大兴安岭、伊春等边缘地区,非遗项目密度低且被低密度区域包围。
|
||||||
|
|
||||||
|
(3)**高-低异常区**:黑河(项目密度较高但周边密度较低),体现了边疆地区的文化特殊性。
|
||||||
|
|
||||||
|
### 3.3 民族特色空间关联
|
||||||
|
|
||||||
|
#### 3.3.1 满族非遗项目分布
|
||||||
|
|
||||||
|
满族非遗项目主要集中在:
|
||||||
|
(1)**宁安地区**:清代宁古塔(今宁安市)是满族龙兴之地,保留了丰富的满族文化遗产,如满族说部、满族剪纸、满族刺绣等。
|
||||||
|
(2)**阿城地区**:金代上京会宁府(今阿城市)是满族先祖女真人的早期都城,保留了满族萨满文化、满族秧歌等。
|
||||||
|
(3)**哈尔滨地区**:作为现代都市,哈尔滨融合了满族、汉族等多种民族文化,如哈尔滨冰雕(融合满族冰雪文化)。
|
||||||
|
|
||||||
|
#### 3.3.2 赫哲族非遗项目分布
|
||||||
|
|
||||||
|
赫哲族非遗项目主要集中在黑龙江、乌苏里江沿岸:
|
||||||
|
(1)**同江市**:赫哲族主要聚居地,保留了大量赫哲族传统文化,如赫哲族鱼皮制作技艺、赫哲族伊玛堪(说唱艺术)。
|
||||||
|
(2)**抚远市**:位于黑龙江、乌苏里江汇合处,是赫哲族渔猎文化的典型代表地。
|
||||||
|
(3)**饶河县**:赫哲族叉草球等传统体育项目的重要传承地。
|
||||||
|
|
||||||
|
#### 3.3.3 鄂伦春族非遗项目分布
|
||||||
|
|
||||||
|
鄂伦春族非遗项目主要分布在大兴安岭地区:
|
||||||
|
(1)**黑河市**:鄂伦春族桦树皮制作技艺、鄂伦春族古伦木沓节(传统节日)。
|
||||||
|
(2)**大兴安岭地区**:鄂伦春族狩猎文化、鄂伦春族兽皮制作技艺。
|
||||||
|
(3)**呼玛县、塔河县**:鄂伦春族摩苏昆(说唱艺术)。
|
||||||
|
|
||||||
|
#### 3.3.4 达斡尔族非遗项目分布
|
||||||
|
|
||||||
|
达斡尔族非遗项目主要分布在齐齐哈尔地区:
|
||||||
|
(1)**梅里斯达斡尔族区**:达斡尔族传统歌舞、达斡尔族哈库麦勒舞(传统舞蹈)。
|
||||||
|
(2)**富拉尔基区**:达斡尔族鲁日格勒舞、达斡尔族民间乐器制作技艺。
|
||||||
|
(3)**讷河市**:达斡尔族传统体育项目。
|
||||||
|
|
||||||
|
#### 3.3.5 民族混合区特征
|
||||||
|
|
||||||
|
在哈尔滨、齐齐哈尔等中心城市,多民族非遗项目共存,形成了民族融合的文化景观。例如:
|
||||||
|
- 哈尔滨既有满族说部,又有朝鲜族农乐舞,还有汉族的东北大鼓。
|
||||||
|
- 齐齐哈尔既有达斡尔族传统歌舞,又有蒙古族四胡音乐,还有鄂温克族驯鹿习俗。
|
||||||
|
|
||||||
|
这种多民族非遗项目共存的格局,体现了黑龙江各民族在长期历史进程中的文化交流与融合。
|
||||||
|
|
||||||
|
### 3.4 类别-环境关联分析
|
||||||
|
|
||||||
|
#### 3.4.1 传统技艺与自然资源
|
||||||
|
|
||||||
|
黑龙江传统技艺类非遗项目与自然资源存在密切关联:
|
||||||
|
|
||||||
|
(1)**森林资源依赖型**:
|
||||||
|
- 桦树皮制作技艺(鄂伦春族):依赖大兴安岭地区的白桦林资源
|
||||||
|
- 木雕技艺(满族、汉族):依赖小兴安岭、张广才岭等森林资源
|
||||||
|
- 柳编技艺(汉族):依赖松嫩平原的柳树资源
|
||||||
|
|
||||||
|
(2)**江河资源依赖型**:
|
||||||
|
- 鱼皮制作技艺(赫哲族):依赖黑龙江、乌苏里江的渔业资源
|
||||||
|
- 皮革制作技艺(鄂伦春族、鄂温克族):依赖狩猎获得的兽皮资源
|
||||||
|
- 网扣制作技艺(汉族):依赖松花江、嫩江的渔业资源
|
||||||
|
|
||||||
|
(3)**农业资源依赖型**:
|
||||||
|
- 剪纸技艺(满族、汉族):依赖农业社会的纸张资源
|
||||||
|
- 刺绣技艺(各民族):依赖丝绸、棉布等纺织品资源
|
||||||
|
- 面食制作技艺(汉族、回族):依赖小麦、玉米等农产品资源
|
||||||
|
|
||||||
|
#### 3.4.2 民俗活动与气候
|
||||||
|
|
||||||
|
黑龙江民俗类非遗项目与寒地气候存在密切关联:
|
||||||
|
|
||||||
|
(1)**冰雪民俗**:
|
||||||
|
- 哈尔滨冰雕:利用严寒气候的冰雪资源
|
||||||
|
- 满族雪地走百病:适应冰雪环境的传统习俗
|
||||||
|
- 鄂伦春族冰雪狩猎:充分利用冬季积雪环境的狩猎活动
|
||||||
|
|
||||||
|
(2)**季节性民俗**:
|
||||||
|
- 鄂伦春族古伦木沓节:春季举行的祭祀活动
|
||||||
|
- 达斡尔族库木勒节:春季采集野菜的传统节日
|
||||||
|
- 蒙古族那达慕:秋季举行的体育竞技活动
|
||||||
|
|
||||||
|
(3)**寒地适应型民俗**:
|
||||||
|
- 鄂温克族驯鹿习俗:适应寒地森林环境的游牧文化
|
||||||
|
- 满族酸菜腌制:适应冬季漫长寒冷的饮食文化
|
||||||
|
- 东北大秧歌:适应冬季室内的娱乐活动
|
||||||
|
|
||||||
|
#### 3.4.3 传统美术与地理
|
||||||
|
|
||||||
|
黑龙江传统美术类非遗项目与地理环境存在关联:
|
||||||
|
|
||||||
|
(1)**东北平原特色**:
|
||||||
|
- 满族剪纸:题材多反映东北平原的农耕生活
|
||||||
|
- 海伦剪纸:融合东北平原的农业文化和满族文化元素
|
||||||
|
|
||||||
|
(2)**沿江地区特色**:
|
||||||
|
- 赫哲族鱼皮画:利用鱼皮材料创作的独特艺术形式
|
||||||
|
- 满族刺绣:融合满族文化和沿江地区的文化特色
|
||||||
|
|
||||||
|
(3)**民族地区特色**:
|
||||||
|
- 朝鲜族跳板、秋千:反映朝鲜族聚居区的文化特色
|
||||||
|
- 蒙古族图案:反映草原民族的艺术风格
|
||||||
|
|
||||||
|
#### 3.4.4 空间自相关性检验
|
||||||
|
|
||||||
|
对非遗类别与环境类型进行空间自相关分析,结果如表7所示。
|
||||||
|
|
||||||
|
**表7 类别-环境空间自相关分析**
|
||||||
|
|
||||||
|
| 类别-环境组合 | Moran's I | p值 | 显著性 |
|
||||||
|
|---------------|-----------|-----|--------|
|
||||||
|
| 传统技艺-森林资源 | 0.421 | <0.01 | 显著正相关 |
|
||||||
|
| 传统技艺-江河资源 | 0.387 | <0.01 | 显著正相关 |
|
||||||
|
| 民俗-寒地气候 | 0.356 | <0.01 | 显著正相关 |
|
||||||
|
| 传统美术-平原地形 | 0.298 | <0.05 | 显著正相关 |
|
||||||
|
| 传统舞蹈-草原环境 | 0.245 | <0.05 | 显著正相关 |
|
||||||
|
|
||||||
|
从表7可以看出,所有组合的Moran's I值均为正值且显著,表明非遗类别与环境类型之间存在显著的空间正相关关系,即特定的非遗类别倾向于在特定的环境中集聚分布。
|
||||||
|
|
||||||
|
### 3.5 典型案例分析
|
||||||
|
|
||||||
|
为深入理解黑龙江非遗的空间关联规律,本研究选取三个典型案例进行深入分析。
|
||||||
|
|
||||||
|
#### 3.5.1 案例一:鄂伦春族桦树皮制作技艺
|
||||||
|
|
||||||
|
**基本信息**:
|
||||||
|
- 项目级别:国家级
|
||||||
|
- 批次:第一批(2006年)
|
||||||
|
- 申报地区:黑龙江省黑河市
|
||||||
|
- 类别:传统技艺
|
||||||
|
|
||||||
|
**空间分布**:
|
||||||
|
主要分布在大兴安岭地区的黑河市、呼玛县、塔河县等地,这些地区拥有丰富的白桦林资源。
|
||||||
|
|
||||||
|
**环境依赖**:
|
||||||
|
- **森林资源**:桦树皮制作技艺直接依赖大兴安岭地区的白桦林资源
|
||||||
|
- **狩猎文化**:鄂伦春族传统上以狩猎为生,桦树皮制品(如桦皮船、桦皮桶)是狩猎生活的重要工具
|
||||||
|
- **气候适应**:桦树皮具有防水、保暖等特性,适应寒地气候
|
||||||
|
|
||||||
|
**民族特征**:
|
||||||
|
- **鄂伦春族特色**:体现了鄂伦春族"依山而居、逐兽而猎"的游猎文化
|
||||||
|
- **口传心授**:技艺传承主要依靠家族传承和师徒传承,缺乏文字记载
|
||||||
|
- **濒危状况**:随着定居化和现代化,传统狩猎文化逐渐衰落,技艺传承面临严峻挑战
|
||||||
|
|
||||||
|
**空间关联**:
|
||||||
|
该项目集中分布在大兴安岭森林文化区,与鄂伦春族其他非遗项目(如古伦木沓节、摩苏昆说唱)形成文化集聚,体现了民族非遗项目与地理环境的高度关联。
|
||||||
|
|
||||||
|
#### 3.5.2 案例二:赫哲族鱼皮制作技艺
|
||||||
|
|
||||||
|
**基本信息**:
|
||||||
|
- 项目级别:国家级
|
||||||
|
- 批次:第一批(2006年)
|
||||||
|
- 申报地区:黑龙江省佳木斯市同江市
|
||||||
|
- 类别:传统技艺
|
||||||
|
|
||||||
|
**空间分布**:
|
||||||
|
主要分布在黑龙江、乌苏里江沿岸的同江市、抚远市、饶河县等地,这些地区拥有丰富的渔业资源。
|
||||||
|
|
||||||
|
**环境依赖**:
|
||||||
|
- **江河资源**:鱼皮制作技艺直接依赖黑龙江、乌苏里江的渔业资源
|
||||||
|
- **渔猎文化**:赫哲族传统上以渔猎为生,鱼皮制品(如鱼皮衣、鱼皮靴)是渔猎生活的重要服饰
|
||||||
|
- **气候适应**:鱼皮具有防水、耐磨、保暖等特性,适应寒地水域环境
|
||||||
|
|
||||||
|
**民族特征**:
|
||||||
|
- **赫哲族特色**:体现了赫哲族"沿江而居、以渔为生"的渔猎文化
|
||||||
|
- **材料创新**:利用鱼类资源创造独特的鱼皮制作技艺,体现了赫哲族的环境适应智慧
|
||||||
|
- **濒危状况**:随着渔业资源减少和现代服饰普及,鱼皮制作技艺面临传承危机
|
||||||
|
|
||||||
|
**空间关联**:
|
||||||
|
该项目集中在黑龙江、乌苏里江沿江文化区,与赫哲族其他非遗项目(如伊玛堪说唱、叉草球体育)形成文化集聚,体现了沿江少数民族非遗项目与江河环境的高度关联。
|
||||||
|
|
||||||
|
#### 3.5.3 案例三:满族说部
|
||||||
|
|
||||||
|
**基本信息**:
|
||||||
|
- 项目级别:国家级
|
||||||
|
- 批次:第一批(2006年)
|
||||||
|
- 申报地区:黑龙江省宁安市、阿城市、哈尔滨市等
|
||||||
|
- 类别:民间文学
|
||||||
|
|
||||||
|
**空间分布**:
|
||||||
|
主要分布在满族聚居的宁安市(清代宁古塔)、阿城市(金代上京会宁府)、哈尔滨市等地区。
|
||||||
|
|
||||||
|
**环境依赖**:
|
||||||
|
- **历史地理**:满族说部与满族的历史迁徙路线密切相关,从长白山麓到松嫩平原
|
||||||
|
- **都市环境**:清代宁古塔、上京会宁府等政治经济中心是满族说部的重要传承地
|
||||||
|
- **文化环境**:满族说部的传承需要稳定的听众群体和文化环境
|
||||||
|
|
||||||
|
**民族特征**:
|
||||||
|
- **满族特色**:满族说部是满族口头传统的集大成者,包含神话、传说、史诗等多种形式
|
||||||
|
- **历史记忆**:满族说部记录了满族从祖先起源到建立清朝的历史进程,是满族历史文化的"活化石"
|
||||||
|
- **濒危状况**:随着现代化和语言环境变化,满族说部的传承面临严峻挑战
|
||||||
|
|
||||||
|
**空间关联**:
|
||||||
|
该项目分布在哈尔滨都市文化区和宁安、阿城等满族历史中心,与满族其他非遗项目(如满族剪纸、满族刺绣)形成文化集聚,体现了满族非遗项目与历史地理环境的密切关联。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 讨论
|
||||||
|
|
||||||
|
### 4.1 方法学创新
|
||||||
|
|
||||||
|
#### 4.1.1 大语言模型在非遗领域的应用优势
|
||||||
|
|
||||||
|
本研究验证了DeepSeek大语言模型在非遗知识抽取中的有效性,主要体现在以下三个方面:
|
||||||
|
|
||||||
|
(1)**减少人工标注成本**。传统知识抽取方法需要大量人工标注的训练数据,成本高昂。周正达等 [@8LEDXI3J] 和陈昱成等 [@NLCHJIZL] 的研究均指出,标注数据的获取是知识图谱构建的瓶颈。本研究利用DeepSeek模型的零样本/少样本学习能力,无需大量标注数据即可实现高质量的知识抽取,大大降低了人工成本。
|
||||||
|
|
||||||
|
(2)**提高知识抽取准确率**。雒伟群等 [@2LWIKRWR] 的研究表明,DeepSeek-R1模型在文物知识抽取任务中F1值达86.25%,较BERT-BiLSTM-CRF模型提升3.13个百分点。本研究的实际测试也表明,DeepSeek模型在非遗知识抽取中的准确率超过90%,显著优于传统方法。
|
||||||
|
|
||||||
|
(3)**支持复杂语义理解**。张卫等 [@2F3PRYB4] 的研究表明,大语言模型通过思维链(CoT)引导,能够进行复杂的语义推理。在大语言模型应用于文化遗产数字化方面,黄刚等 [@XIVILIDP] 提出了利用大语言模型对中国家谱进行结构化处理和知识图谱构建的方法,抽取了约30万条结构化记录,验证了大语言模型在文化遗产领域的有效性。本研究设计的多阶段提示策略,引导DeepSeek模型逐步完成实体识别、关系抽取和属性填充,实现了对非遗文本的深度语义理解。
|
||||||
|
|
||||||
|
#### 4.1.2 本地化本体设计的价值
|
||||||
|
|
||||||
|
本研究设计的本体模型在复用CIDOC CRM的基础上,增加了民族特色和自然环境两个本地化概念类型,体现了以下价值:
|
||||||
|
|
||||||
|
(1)**平衡通用性与特殊性**。张萌萌等 [@NNJSH2A9] 指出,地方非遗知识图谱需要在通用模型的基础上体现地域特色。本研究通过扩展民族特色和自然环境概念,既保证了与其他知识图谱的互操作性,又充分体现了黑龙江非遗的独特性。
|
||||||
|
|
||||||
|
(2)**增强语义表达能力**。陈涛等 [@YKHCTW2K] 强调,文化遗产知识表征需要兼顾多模态和地域特色。本研究增加的民族特色和自然环境概念,能够更准确地描述黑龙江非遗与民族、环境之间的关联关系。
|
||||||
|
|
||||||
|
(3)**支持空间分析**。陆锋等 [@6JYY2GXD] 指出,地理知识图谱需要强调时间和空间特征。本研究增加的自然环境概念,为非遗空间分析提供了语义基础,能够揭示非遗项目与地理环境之间的关联规律。
|
||||||
|
|
||||||
|
#### 4.1.3 多维度校验机制的必要性
|
||||||
|
|
||||||
|
雒伟群等 [@2LWIKRWR] 提出的多维度校验机制在本研究中得到了验证。本研究从逻辑一致性、领域规范性和事实准确性三个方面对DeepSeek抽取的知识进行校验,确保了知识图谱的质量:
|
||||||
|
|
||||||
|
(1)**逻辑一致性检查**能够发现时间顺序错误(如传承人出生年份晚于项目批准年份)、关系冲突(如一个人同时属于多个不兼容的民族)等问题。
|
||||||
|
|
||||||
|
(2)**领域规范性检查**能够确保项目类别属于国家标准的10大类之一、民族名称属于黑龙江世居民族等。
|
||||||
|
|
||||||
|
(3)**事实准确性检查**通过多源数据验证,能够纠正实体名称错误、地名拼写错误等事实性问题。
|
||||||
|
|
||||||
|
### 4.2 空间分布规律的发现
|
||||||
|
|
||||||
|
#### 4.2.1 "核心-边缘"分布模式
|
||||||
|
|
||||||
|
本研究揭示了黑龙江非遗的"核心-边缘"空间分布模式:
|
||||||
|
|
||||||
|
(1)**核心区域**:哈尔滨、齐齐哈尔、牡丹江等中心城市,非遗项目密度高、类别丰富、多民族融合。这些地区的非遗项目具有以下特点:
|
||||||
|
- 数量多:哈尔滨45项、齐齐哈尔38项、牡丹江32项,占全省的46.5%
|
||||||
|
- 类别全:涵盖10大类别,传统技艺、民俗、传统美术等类别齐全
|
||||||
|
- 民族多:满族、汉族、朝鲜族、蒙古族等多民族非遗项目共存
|
||||||
|
- 创新性强:传统与现代结合,如哈尔滨冰雕、满族剪纸等
|
||||||
|
|
||||||
|
(2)**边缘区域**:大兴安岭、伊春等边疆地区,非遗项目数量不多但具有独特的民族文化特色。这些地区的非遗项目具有以下特点:
|
||||||
|
- 数量少:大兴安岭2项、伊春5项
|
||||||
|
- 特色鲜明:鄂伦春族、鄂温克族等少数民族非遗项目集中分布
|
||||||
|
- 环境依赖性强:与森林、江河等自然环境高度关联
|
||||||
|
- 濒危程度高:受现代化冲击大,传承面临严峻挑战
|
||||||
|
|
||||||
|
这种"核心-边缘"分布模式与张雪英等 [@UCWTRX2J] 提出的地理知识分布规律一致,即核心区域集聚度高、边缘区域集聚度低。
|
||||||
|
|
||||||
|
#### 4.2.2 文化生态区的识别
|
||||||
|
|
||||||
|
本研究通过聚类分析识别出三大文化生态区:
|
||||||
|
|
||||||
|
(1)**森林文化区**(大兴安岭地区):
|
||||||
|
- 空间范围:黑河、大兴安岭、伊春北部
|
||||||
|
- 环境特征:森林资源丰富,气候寒冷
|
||||||
|
- 民族特色:鄂伦春族、鄂温克族狩猎文化
|
||||||
|
- 代表项目:桦树皮制作技艺、兽皮制作技艺、摩苏昆说唱
|
||||||
|
- 文化特征:体现狩猎民族与森林环境的和谐共生
|
||||||
|
|
||||||
|
(2)**农耕文化区**(松嫩平原):
|
||||||
|
- 空间范围:哈尔滨、齐齐哈尔、绥化、大庆
|
||||||
|
- 环境特征:平原地形,农业资源丰富
|
||||||
|
- 民族特色:满族、汉族农耕文化
|
||||||
|
- 代表项目:满族说部、满族剪纸、海伦剪纸、东北大秧歌
|
||||||
|
- 文化特征:体现农耕民族的定居文化和民俗传统
|
||||||
|
|
||||||
|
(3)**江河文化区**(黑龙江、乌苏里江沿岸):
|
||||||
|
- 空间范围:佳木斯、牡丹江、双鸭山、鹤岗
|
||||||
|
- 环境特征:江河资源丰富,沿江地理环境
|
||||||
|
- 民族特色:赫哲族、满族、朝鲜族渔猎文化
|
||||||
|
- 代表项目:鱼皮制作技艺、伊玛堪说唱、满族刺绣、朝鲜族农乐舞
|
||||||
|
- 文化特征:体现沿江民族的渔猎文化和跨境文化交流
|
||||||
|
|
||||||
|
这三大文化生态区的识别,与陆锋等 [@6JYY2GXD] 提出的地理知识表达模型一致,即地理知识需要考虑"地理概念–地理实体–地理关系"三个层次。
|
||||||
|
|
||||||
|
#### 4.2.3 民族非遗项目的空间隔离与融合
|
||||||
|
|
||||||
|
本研究发现了民族非遗项目的空间隔离与融合并存的现象:
|
||||||
|
|
||||||
|
(1)**空间隔离**:
|
||||||
|
- 鄂伦春族非遗项目集中在大兴安岭地区,与其他民族项目空间距离较远
|
||||||
|
- 赫哲族非遗项目集中在黑龙江、乌苏里江沿岸,形成独特的沿江文化带
|
||||||
|
- 蒙古族非遗项目集中在大庆草原地区,与农耕文化区相对独立
|
||||||
|
|
||||||
|
这种空间隔离现象,体现了地理环境对文化形成的影响。张雪英等 [@UCWTRX2J] 指出,地理知识具有特定的时空特征,不同地理环境孕育不同的文化形态。
|
||||||
|
|
||||||
|
(2)**空间融合**:
|
||||||
|
- 哈尔滨作为核心都市,融合了满族、汉族、朝鲜族等多种民族文化
|
||||||
|
- 齐齐哈尔融合了达斡尔族、蒙古族、鄂温克族等民族文化
|
||||||
|
- 牡丹江融合了满族、朝鲜族、汉族等民族文化
|
||||||
|
|
||||||
|
这种空间融合现象,体现了多民族聚居区的文化交流与融合。范炜等 [@YUBSWZ5V] 指出,AI时代的文化遗产活化利用需要关注文化融合与创新。
|
||||||
|
|
||||||
|
### 4.3 环境依赖性与文化适应
|
||||||
|
|
||||||
|
#### 4.3.1 自然环境对非遗形成的影响
|
||||||
|
|
||||||
|
本研究揭示了自然环境对非遗形成的深刻影响:
|
||||||
|
|
||||||
|
(1)**材料依赖**:
|
||||||
|
- 桦树皮制作技艺直接依赖白桦林资源
|
||||||
|
- 鱼皮制作技艺直接依赖江河渔业资源
|
||||||
|
- 兽皮制作技艺直接依赖狩猎获得的动物皮毛资源
|
||||||
|
|
||||||
|
这种材料依赖体现了非遗项目的环境决定性。赵万青等 [@JH5C6HKD] 指出,文化遗产的多模态表征需要考虑材料、工艺、环境等多重要素。
|
||||||
|
|
||||||
|
(2)**气候适应**:
|
||||||
|
- 冰雪民俗(如哈尔滨冰雕、满族雪地走百病)是对寒地气候的直接适应
|
||||||
|
- 鄂温克族驯鹿习俗是对高寒森林环境的适应
|
||||||
|
- 东北大秧歌等室内娱乐活动是对冬季漫长气候的适应
|
||||||
|
|
||||||
|
这种气候适应体现了非遗项目的环境适应性。陆锋等 [@6JYY2GXD] 指出,地理知识需要考虑时空特征和环境依赖性。
|
||||||
|
|
||||||
|
(3)**地形影响**:
|
||||||
|
- 山地狩猎文化(鄂伦春族、鄂温克族)依赖山地地形
|
||||||
|
- 平原农耕文化(满族、汉族)依赖平原地形
|
||||||
|
- 江河渔猎文化(赫哲族)依赖江河地形
|
||||||
|
|
||||||
|
这种地形影响体现了非遗项目的地理分异规律。陈涛等 [@YKHCTW2K] 指出,文化遗产知识表征需要扩展空间方位属性,以描述地理环境影响。
|
||||||
|
|
||||||
|
#### 4.3.2 文化适应策略
|
||||||
|
|
||||||
|
本研究发现,黑龙江非遗项目在长期历史进程中形成了多种文化适应策略:
|
||||||
|
|
||||||
|
(1)**技艺创新**:
|
||||||
|
- 材料替代:随着森林资源减少,桦树皮制作技艺逐渐转向艺术化、展示化
|
||||||
|
- 工具改进:传统手工工具逐渐被现代工具替代,提高制作效率
|
||||||
|
- 功能转型:从实用工具向艺术作品、文化纪念品转型
|
||||||
|
|
||||||
|
(2)**功能转型**:
|
||||||
|
- 从实用到艺术:桦树皮船从狩猎工具转向艺术装饰品
|
||||||
|
- 从生产到表演:传统歌舞从生产生活场景转向舞台表演
|
||||||
|
- 从生活到展示:民俗活动从日常生活转向节庆展示
|
||||||
|
|
||||||
|
(3)**传承方式变化**:
|
||||||
|
- 家族传承→学校教育:非遗传承逐渐纳入学校教育体系
|
||||||
|
- 师徒传承→社会传承:非遗传承逐渐社会化、公开化
|
||||||
|
- 口传心授→数字化保护:利用数字技术记录、保存、传播非遗
|
||||||
|
|
||||||
|
魏立才 [@UFAR2JX6] 指出,多模态大模型能够为文化遗产的保护与传承提供新的路径,包括技艺创新、功能转型和传承方式变化。
|
||||||
|
|
||||||
|
### 4.4 保护与传承策略建议
|
||||||
|
|
||||||
|
基于上述发现,本研究提出以下保护与传承策略建议:
|
||||||
|
|
||||||
|
#### 4.4.1 基于空间分布的差异化保护策略
|
||||||
|
|
||||||
|
(1)**核心区域**(哈尔滨、齐齐哈尔等):
|
||||||
|
- 建立非遗保护示范区,整合各类非遗资源
|
||||||
|
- 发展非遗旅游,促进非遗与现代生活的融合
|
||||||
|
- 建设非遗体验馆,提供沉浸式文化体验
|
||||||
|
- 利用知识图谱技术,开发智能导览、文化推荐等服务
|
||||||
|
|
||||||
|
(2)**边缘区域**(大兴安岭、伊春等):
|
||||||
|
- 加强传承人扶持,提供资金和技术支持
|
||||||
|
- 建立文化生态保护区,整体保护民族非遗及其生存环境
|
||||||
|
- 开展非遗数字化保护,利用3D建模、VR等技术永久保存非遗项目
|
||||||
|
- 鼓励传承人收徒授艺,扩大传承人群
|
||||||
|
|
||||||
|
(3)**民族区域**(鄂伦春族、赫哲族、达斡尔族等聚居区):
|
||||||
|
- 保护文化生态完整性,避免非遗项目脱离其生存环境
|
||||||
|
- 开展双语教育,保护和传承民族语言
|
||||||
|
- 支持民族节庆活动,为非遗传承提供实践场景
|
||||||
|
- 建立民族非遗数据库,系统记录民族非遗项目
|
||||||
|
|
||||||
|
#### 4.4.2 数字化保护路径
|
||||||
|
|
||||||
|
(1)**知识图谱应用**:
|
||||||
|
- 开发智能检索系统,支持语义搜索和关联推荐
|
||||||
|
- 构建问答系统,提供非遗知识问答服务 [@EC26X7PS; @FMHA6MVH]
|
||||||
|
- 可视化展示知识图谱,帮助公众理解非遗项目之间的关联
|
||||||
|
- 利用知识图谱进行空间分析,支持文化空间规划
|
||||||
|
|
||||||
|
(2)**多模态资源整合** [@YKHCTW2K; @JH5C6HKD]:
|
||||||
|
- 采集非遗项目的图像、音视频等多模态资源
|
||||||
|
- 利用多模态大模型进行内容理解和生成
|
||||||
|
- 开发沉浸式VR/AR体验,提供身临其境的文化体验
|
||||||
|
- 构建多模态知识库,支持跨模态检索和推理
|
||||||
|
|
||||||
|
(3)**虚拟现实展示** [@UFAR2JX6]:
|
||||||
|
- 开发非遗VR体验系统,重现非遗项目的历史场景
|
||||||
|
- 利用AR技术,在现实环境中叠加非遗信息
|
||||||
|
- 开发非遗数字博物馆,突破时空限制展示非遗
|
||||||
|
- 构建非遗元宇宙,提供虚拟社交和互动体验
|
||||||
|
|
||||||
|
#### 4.4.3 政策建议
|
||||||
|
|
||||||
|
(1)**文化空间规划**:
|
||||||
|
- 将非遗保护纳入城乡建设规划,保护非遗赖以生存的文化空间
|
||||||
|
- 建立非遗保护利用设施,如非遗展示馆、传承基地等
|
||||||
|
- 规划非遗旅游线路,促进非遗与旅游融合发展
|
||||||
|
- 保护非遗的自然环境,避免非遗项目脱离其生存环境
|
||||||
|
|
||||||
|
(2)**跨区域协作**:
|
||||||
|
- 建立东北三省非遗保护联盟,促进区域交流与合作
|
||||||
|
- 开展跨境非遗保护合作,保护赫哲族等跨境民族的非遗
|
||||||
|
- 建立非遗保护信息共享平台,促进经验交流
|
||||||
|
- 开展联合申报世界非物质文化遗产,提升国际影响力
|
||||||
|
|
||||||
|
(3)**传承人培养**:
|
||||||
|
- 将非遗纳入中小学教育,开展非遗传承普及教育
|
||||||
|
- 支持高校设立非遗相关专业,培养专业人才
|
||||||
|
- 建立传承人津贴制度,保障传承人基本生活
|
||||||
|
- 鼓励传承人收徒授艺,扩大传承人群
|
||||||
|
|
||||||
|
### 4.5 研究局限性
|
||||||
|
|
||||||
|
本研究存在以下局限性:
|
||||||
|
|
||||||
|
(1)**数据完整性**:部分非遗项目缺少详细的地理信息和传承人信息,影响空间分析的精度。未来需要补充采集这些信息。
|
||||||
|
|
||||||
|
(2)**时态信息**:本研究主要关注非遗项目的当前状态,对历史演变过程分析不足。未来需要构建时序知识图谱,追溯非遗项目的历史变迁。
|
||||||
|
|
||||||
|
(3)**关系推理**:部分隐性关系依赖人工标注,自动化程度有待提高。未来需要引入更先进的推理算法,提高关系推理的自动化水平。
|
||||||
|
|
||||||
|
(4)**可视化限制**:当前的知识图谱可视化系统主要展示静态关联,难以展现动态演变过程。未来需要开发时序可视化、动态演化等功能。
|
||||||
|
|
||||||
|
### 4.6 未来研究方向
|
||||||
|
|
||||||
|
基于本研究的发现和局限,未来研究可以从以下方向展开:
|
||||||
|
|
||||||
|
(1)**时序知识图谱**:构建非遗历史演变知识图谱,追溯非遗项目从起源到现在的历史变迁,分析其演化规律和影响因素。
|
||||||
|
|
||||||
|
(2)**多模态融合**:整合非遗项目的文本、图像、音视频等多模态资源,构建多模态知识图谱,支持跨模态检索和推理 [@JH5C6HKD]。
|
||||||
|
|
||||||
|
(3)**跨区域对比**:将黑龙江非遗与吉林、辽宁等东北其他省份进行对比研究,揭示东北非遗的共性与差异,构建东北区域非遗知识图谱。
|
||||||
|
|
||||||
|
(4)**智能应用**:基于知识图谱开发智能问答系统 [@EC26X7PS]、文化推荐系统、非遗传承预警系统等,为非遗保护提供智能化工具。
|
||||||
|
|
||||||
|
(5)**预测分析**:基于历史数据和当前趋势,预测非遗项目的传承发展趋势,识别高风险项目,为保护决策提供科学依据。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 结论
|
||||||
|
|
||||||
|
### 5.1 主要发现总结
|
||||||
|
|
||||||
|
本研究针对黑龙江省非物质文化遗产知识组织与空间分析的问题,提出了基于DeepSeek大语言模型的非遗知识图谱构建方法,并系统分析了其空间分布特征与关联性。主要发现如下:
|
||||||
|
|
||||||
|
(1)**成功构建了黑龙江省非遗知识图谱**。通过DeepSeek模型的知识抽取,构建了包含639个节点、639条关系的知识图谱,涵盖12类核心概念、40余种关系类型。DeepSeek模型在知识抽取中的成功率达100%,平均处理时间3.6秒/节点,实体识别准确率95.2%,关系抽取准确率91.7%。
|
||||||
|
|
||||||
|
(2)**验证了DeepSeek大模型在非遗知识抽取中的有效性**。与传统方法相比,本研究方法具有明显优势:减少人工标注成本、提高知识抽取准确率、支持复杂语义理解。多维度校验机制确保了知识图谱的质量,逻辑一致性达98.3%。
|
||||||
|
|
||||||
|
(3)**揭示了黑龙江非遗"核心-边缘"的空间分布模式**。哈尔滨、齐齐哈尔等中心城市是非遗分布的核心区域,项目密度高、类别丰富、多民族融合;大兴安岭、伊春等边缘地区项目数量不多但具有独特的民族文化特色。
|
||||||
|
|
||||||
|
(4)**识别了三大文化生态区**:森林文化区(大兴安岭地区,鄂伦春族、鄂温克族狩猎文化)、农耕文化区(松嫩平原,满族、汉族农耕文化)、江河文化区(黑龙江、乌苏里江沿岸,赫哲族、满族、朝鲜族渔猎文化)。
|
||||||
|
|
||||||
|
(5)**发现了非遗项目与民族特色、地理环境的显著空间关联**。Moran's I指数为0.356(p<0.01),表明非遗项目存在显著的空间正自相关。传统技艺-森林资源(Moran's I=0.421)、传统技艺-江河资源(Moran's I=0.387)、民俗-寒地气候(Moran's I=0.356)等组合均表现出显著的空间正相关。
|
||||||
|
|
||||||
|
### 5.2 理论贡献
|
||||||
|
|
||||||
|
本研究的理论贡献主要体现在以下三个方面:
|
||||||
|
|
||||||
|
(1)**提出了基于大语言模型的非遗知识图谱构建方法**。本研究设计了多阶段任务分解的联合抽取策略,包括实体属性提取、三元组关系抽取和多维度校验,为非遗知识图谱构建提供了新方法。
|
||||||
|
|
||||||
|
(2)**设计了融合地域特色的本体模型**。本研究在复用CIDOC CRM的基础上,增加了民族特色和自然环境两个本地化概念类型,为区域性非遗知识组织提供了可复制的范式。
|
||||||
|
|
||||||
|
(3)**建立了非遗空间关联分析框架**。本研究综合运用核密度估计、聚类分析、G统计量、空间自相关分析等方法,为非遗空间分析提供了系统的方法论框架。
|
||||||
|
|
||||||
|
### 5.3 实践价值
|
||||||
|
|
||||||
|
本研究的实践价值主要体现在以下三个方面:
|
||||||
|
|
||||||
|
(1)**为黑龙江非遗数字化保护提供技术工具**。构建的知识图谱和可视化系统可支持非遗资源的智能检索、知识问答和空间分析,为非遗数字化保护提供了技术支撑。
|
||||||
|
|
||||||
|
(2)**为文化空间规划提供科学依据**。揭示的"核心-边缘"分布模式和三大文化生态区,为文化空间规划、保护政策制定提供了科学依据。
|
||||||
|
|
||||||
|
(3)**促进非遗文化的传播与教育应用**。构建的知识图谱可应用于非遗教育、文化旅游、文化创意等领域,促进非遗文化的传播与创新性发展。
|
||||||
|
|
||||||
|
### 5.4 展望
|
||||||
|
|
||||||
|
随着"大模型+知识图谱"双轮驱动范式的发展 [@LJ5QBY4I],非遗数字化保护将进入智能化、个性化、沉浸化的新阶段。未来应进一步探索以下方向:
|
||||||
|
|
||||||
|
(1)**多模态知识表征**:整合文本、图像、音视频等多模态资源,构建多模态非遗知识图谱 [@YKHCTW2K; @JH5C6HKD]。
|
||||||
|
|
||||||
|
(2)**智能问答系统**:基于知识图谱和大模型,开发非遗智能问答系统,提供精准的知识问答服务 [@EC26X7PS; @FMHA6MVH]。
|
||||||
|
|
||||||
|
(3)**文化遗产虚拟现实**:利用VR/AR技术,开发沉浸式非遗体验系统,提供身临其境的文化体验 [@UFAR2JX6]。
|
||||||
|
|
||||||
|
(4)**预测与决策支持**:基于知识图谱和时空分析,开发非遗传承预警系统和保护决策支持系统,为非遗保护提供智能化工具。
|
||||||
|
|
||||||
|
(5)**跨区域协同保护**:构建东北区域非遗知识图谱,开展跨区域协同保护,提升非遗保护的整体效能。
|
||||||
|
|
||||||
|
通过上述方向的研究与实践,将为非遗保护与传承提供更强有力的技术支撑,推动中华优秀传统文化创造性转化、创新性发展。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 参考文献
|
||||||
|
|
||||||
|
陈涛, 张欣, 冯卓彤, 杨鑫. 文化遗产多模态资源知识统一表征模型构建研究 [@YKHCTW2K]. 中国图书馆学报, 2025, 51(6).
|
||||||
|
|
||||||
|
陈涛, 刘炜, 单蓉蓉, 朱庆华. 知识图谱在数字人文中的应用研究 [@72HK557Y]. 中国图书馆学报, 2019, 45(6), 34-49.
|
||||||
|
|
||||||
|
陈昱成, 黎洋, 刘江峰, 杨帆. AIGC视角下非物质文化遗产知识图谱的构建研究 [@NLCHJIZL]. 科技情报研究, 2024, 6(2).
|
||||||
|
|
||||||
|
范炜, 曾蕾. AI新时代面向文化遗产活化利用的智慧数据生成路径探析 [@YUBSWZ5V]. 中国图书馆学报, 2024, 50(2).
|
||||||
|
|
||||||
|
敖若瑶. 文化遗产领域知识图谱发展趋势与前沿进展研究 [@FYB5RXLC]. 科技与创新, 2025(17).
|
||||||
|
|
||||||
|
李嘉仪, 李想, 马小柯, 胡浩天, 王丽华. 文理融通:AGI时代的数字人文——第六届中国数字人文年会(CDH2024)会议综述 [@RPNDCWWB]. 数字人文研究, 2025, 5(1).
|
||||||
|
|
||||||
|
李根. 基于大模型技术的档案文化遗产自动问答平台构建研究 [@EC26X7PS]. 山西档案, 2024(9).
|
||||||
|
|
||||||
|
刘文俏. 大模型与古籍档案文化遗产数字化:价值、挑战与应对 [@N9VGUZD6]. 山西档案, 2024(1).
|
||||||
|
|
||||||
|
刘彦超, 刘键, 席上琳, 晁溪蕊, 侯娜, 朱文莲. 基于Neo4j的中轴线艺术价值数字化知识图谱研究 [@49Q5D2DH]. 包装工程, 2024, 45(8).
|
||||||
|
|
||||||
|
陆锋, 余丽, 仇培元. 论地理知识图谱 [@6JYY2GXD]. 地球信息科学学报, 2017, 19(6), 723-734.
|
||||||
|
|
||||||
|
彭纪扬, 郑昂. 基于Neo4j的湘西地区旅游知识图谱构建研究 [@5VBJ427S]. 科技资讯, 2025, 23(8).
|
||||||
|
|
||||||
|
雒伟群, 刘华瑞. 基于大语言模型的唐蕃古道文物知识图谱构建研究 [@2LWIKRWR]. 计算机科学与探索, 2026, 20(3), 801.
|
||||||
|
|
||||||
|
徐怀钰, 赵俊伟, 彭潇, 黄梅荣. 基于大模型的非遗知识图谱与智慧问答系统构建研究 [@FMHA6MVH]. 华东科技, 2025(6).
|
||||||
|
|
||||||
|
杨萌, 张云中, 赵程程. "大模型+知识图谱"双轮驱动的公共数字文化资源管理新范式 [@LJ5QBY4I]. 情报科学, 2025, 43(9).
|
||||||
|
|
||||||
|
张卫, 高鑫, 张予歌. 大语言模型强化学习驱动的文化遗迹叙事文本语义组织方法研究 [@2F3PRYB4]. 图书情报工作, 2025.
|
||||||
|
|
||||||
|
岳丽欣, 刘文云. 国内外领域本体构建方法的比较研究 [@VUF3ZSSW]. 情报理论与实践, 2016, 39(8), 119-125.
|
||||||
|
|
||||||
|
韩牧哲, 高劲松, 李钰. 面向考古类型学的出土陶器器形的知识表示与语义关联构建 [@XTQRLFPW]. 图书情报工作, 2022, 66(12), 92-107.
|
||||||
|
|
||||||
|
张萌萌, 张矛矛. 地方非物质文化遗产知识图谱构建及其思政教育应用 [@NNJSH2A9]. 情报科学, 2025, 43(9).
|
||||||
|
|
||||||
|
张雪英, 张春菊, 吴明光, 闾国年. 顾及时空特征的地理知识图谱构建方法 [@UCWTRX2J]. 中国科学:信息科学, 2020, 50(7), 1019-1032.
|
||||||
|
|
||||||
|
曾子明, 周知, 蒋琳. 基于关联数据的数字人文视觉资源知识组织研究 [@MLPW6LLL]. 情报资料工作, 2018(6), 6-12.
|
||||||
|
|
||||||
|
赵万青, 徐朝阳, 谢智伟, 张少博, 张晓丹, 彭进业. "博古问津":知识图谱增强的文化遗产领域多模态大模型 [@JH5C6HKD]. 西北大学学报(自然科学版), 2025, 55(6).
|
||||||
|
|
||||||
|
周莉娜, 洪亮, 高子阳. 唐诗知识图谱的构建及其智能知识服务设计 [@B79P45VU]. 图书情报工作, 2019, 63(2), 24-33.
|
||||||
|
|
||||||
|
周正达, 王昊, 汪琳, 李晓敏, 周抒, 姚天辰. ChatKG:一种基于大语言模型和提示工程的非遗知识图谱构建框架——以中国非遗陶瓷制作工艺为例 [@8LEDXI3J]. 图书馆杂志, 2025-02-24.
|
||||||
|
|
||||||
|
朱丽雅, 张珺, 洪亮, 罗绍辉, 兰度. 数字人文领域的知识图谱:研究进展与未来趋势 [@LDUKQAWH]. 知识管理论坛, 2022, 7(1), 87-100.
|
||||||
|
|
||||||
|
蒋金亮, 徐云翼, 杨晗, 刘志超. 基于知识图谱和大模型的文化遗产展示和查询方法研究——以大运河文化遗产为例 [@4GYKXZJ5]. 中国名城, 2024, 38(12).
|
||||||
|
|
||||||
|
王左戎, 邓三鸿, 胡畔, 翟姗姗. 价值共创视域下中国传统戏曲知识图谱模式层构建及应用研究 [@IBVAJNMD]. 情报科学, 2025, 43(3), 165-175.
|
||||||
|
|
||||||
|
魏立才. 叙事、认同、沉浸:多模态大模型赋能新时期文化遗产保护与传承的推进策略 [@UFAR2JX6]. 云南民族大学学报(哲学社会科学版), 2025, 42(1).
|
||||||
|
|
||||||
|
黄刚, 林磊. 大语言模型赋能家谱数字化与知识图谱构建研究——以宁波市天一阁博物院实践为例 [@XIVILIDP]. 文化创新比较研究, 2025, 9(24), 189-194.
|
||||||
|
|
||||||
|
汪琳, 王昊, 李晓敏, 邓三鸿. 融合学习扩展的非遗陶瓷工艺领域术语库构建及应用 [@GCLXKNED]. 图书馆论坛, 2024, 44(2), 66-78.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**致谢**
|
||||||
|
|
||||||
|
本研究得到了XXX基金(项目编号:XXX)的资助。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**作者简介**
|
||||||
|
|
||||||
|
第一作者:彭晓,男,哈尔滨工业大学建筑与设计学院副研究员,研究方向为空间智能、数字人文和生态设计。
|
||||||
|
|
||||||
|
通信作者:XXX,性别,职称,学历,Email: xxx@xxx.com,研究方向为文化遗产数字化保护。
|
||||||
File diff suppressed because it is too large
Load Diff
Binary file not shown.
Reference in New Issue
Block a user