Elasticsearch 使用bulk批量导入数据

2024-04-12 11:06:44

批量导入可以合并多个操作，比如index,delete,update,create等等。也可以帮助从一个索引导入到另一个索引。

语法大致如下；

action_and_meta_data\n
optional_source\n
action_and_meta_data\n
optional_source\n
....
action_and_meta_data\n
optional_source\n

需要注意的是，每一条数据都由两行构成（delete除外），其他的命令比如index和create都是由元信息行和数据行组成，update比较特殊它的数据行可能是doc也可能是upsert或者script,如果不了解的朋友可以参考前面的update的翻译。

注意，每一行都是通过\n回车符来判断结束，因此如果你自己定义了json，千万不要使用回车符。不然_bulk命令会报错的！

一个小例子

比如我们现在有这样一个文件，data.json：

{ "index" : { "_index" : "test", "_type" : "type1", "_id" : "1" } }
{ "field1" : "value1" }
{ "index" : { "_index" : "test", "_type" : "_doc" } }
{ "field1" : "value1" }
{ "index":{ "_index": "test2", "_type": "_doc" }}
{"field1":"xxx","field2":"xxxx","field3":"xxx"}

它的第一行定义了_index，_type，_id等信息；第二行定义了字段的信息。

然后执行命令：

curl -XPOST localhost:9200/_bulk --data-binary @data.json

　如果需要指定用户名和密码如下：

curl -XPOST -uuser:password localhost:9200/_bulk --data-binary @data.json

　　就可以看到已经导入进去数据了。

对于其他的index,delete,create,update等操作也可以参考下面的格式：

{ "index" : { "_index" : "test", "_type" : "type1", "_id" : "1" } }
{ "field1" : "value1" }
{ "delete" : { "_index" : "test", "_type" : "type1", "_id" : "2" } }
{ "create" : { "_index" : "test", "_type" : "type1", "_id" : "3" } }
{ "field1" : "value3" }
{ "update" : {"_id" : "1", "_type" : "type1", "_index" : "index1"} }
{ "doc" : {"field2" : "value2"} }