sphinx 同时使用多个索引进行检索探究

2014年2月15日 11:24:34

结论:

1.一次性使用多个索引进行查询的时候,返回的结果集中的fields字段没有什么清楚的意义(也没有找到文档对它的说明)

2.如果程序中一次搜索使用了多个索引,如果它们配置文件中过滤用的属性(aql_attr_uint,sql_field_string...)不全相同,那么最终返回的结果集中,只包含这几个索引*有的属性

实验:

建立两个索引:goods_brand,  goods_cate, 分别是商品信息+品牌信息,商品信息+分类信息

  sql_query = select gid, gid as goodsid, siteid, catename from v_goods_info_cate
sql_attr_uint = siteid
sql_attr_uint = goodsid
sql_field_string = catename ####################### sql_query = select gid, gid as goodsid, siteid, brandname from v_goods_info_brand
sql_attr_uint = siteid
sql_attr_uint = goodsid
sql_field_string = brandname

注:

1. brandname 是商品的品牌名字, catename是商品的分类名字

2. brandname, catename 在索引时,既作为全文索引,又作为属性值返回

3. siteid在两个索引中都有,brandname和catename只在各自的索引中存在

测试程序代码

 $sphObj->AddQuery($keyword, 'goods_brand');
$sphObj->AddQuery($keyword, 'goods_cate');
$sphObj->AddQuery($keyword, 'goods_cate, goods_brand');
$sphObj->AddQuery($keyword, 'goods_brand,goods_cate'); var_dump($rs[0]['fields'], $rs[0]['words'], $rs[0]['matches']);

注:

在程序中做控制:搜索"机"这个字,在goods_cate和goods_brand索引中各只有两条记录符合要求(一共有4条记录):

1.分别执行测试代码的第1行和第2行,并用第6行打印出结果:

 //goods_brand
array (size=1)
0 => string 'brandname' (length=9) array (size=1)
'机' =>
array (size=2)
'docs' => string '10049' (length=5)
'hits' => string '10049' (length=5) array (size=2)
0 =>
array (size=3)
'id' => string '157978' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=3)
'goodsid' => string '157978' (length=6)
'siteid' => string '102' (length=3)
'brandname' => string '无锡一机' (length=12)
1 =>
array (size=3)
'id' => string '157980' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=3)
'goodsid' => string '157980' (length=6)
'siteid' => string '102' (length=3)
'brandname' => string '无锡一机' (length=12) //goods_cate
array (size=1)
0 => string 'catename' (length=8) array (size=1)
'机' =>
array (size=2)
'docs' => string '43986' (length=5)
'hits' => string '43986' (length=5) array (size=2)
0 =>
array (size=3)
'id' => string '158010' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=3)
'goodsid' => string '158010' (length=6)
'siteid' => string '102' (length=3)
'catename' => string '磨齿机' (length=9)
1 =>
array (size=3)
'id' => string '158014' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=3)
'goodsid' => string '158014' (length=6)
'siteid' => string '102' (length=3)
'catename' => string '旋压机' (length=9)

注:

每个索引单独被使用时,各对应两条记录(一共有4条记录)

每条匹配记录中的'attrs'中有siteid+brandname,或者,siteid+catename

2.当用一次查询用多个索引时:分别执行第3行和第4行,并用第6行打印出结果:

 //goods_brand在前,goods_cate在后
array (size=1)
0 => string 'brandname' (length=9) array (size=1)
'机' =>
array (size=2)
'docs' => string '54035' (length=5)
'hits' => string '54035' (length=5) array (size=4)
0 =>
array (size=3)
'id' => string '157978' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '157978' (length=6)
'siteid' => string '102' (length=3)
1 =>
array (size=3)
'id' => string '157980' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '157980' (length=6)
'siteid' => string '102' (length=3)
2 =>
array (size=3)
'id' => string '158010' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '158010' (length=6)
'siteid' => string '102' (length=3)
3 =>
array (size=3)
'id' => string '158014' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '158014' (length=6)
'siteid' => string '102' (length=3) //goods_cate在前,goods_brand在后
array (size=1)
0 => string 'catename' (length=8) array (size=1)
'机' =>
array (size=2)
'docs' => string '54035' (length=5)
'hits' => string '54035' (length=5) array (size=4)
0 =>
array (size=3)
'id' => string '157978' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '157978' (length=6)
'siteid' => string '102' (length=3)
1 =>
array (size=3)
'id' => string '157980' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '157980' (length=6)
'siteid' => string '102' (length=3)
2 =>
array (size=3)
'id' => string '158010' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '158010' (length=6)
'siteid' => string '102' (length=3)
3 =>
array (size=3)
'id' => string '158014' (length=6)
'weight' => string '1' (length=1)
'attrs' =>
array (size=2)
'goodsid' => string '158014' (length=6)
'siteid' => string '102' (length=3)

注:

两个索引被同时使用,只有先后顺序不一样时,4条记录都得到了(这样的结果是对的)

但是第3行和第47行的代码键值对表明,返回的结果集中的fields值没有什么特别的含义(至少我不知到,难道只和排在前边的索引使用的全文索引字段同步?肯定有什么意义,只是我没有总结到吧)

另外,查看结果知道,每一条匹配记录的'attrs'数组中只有siteid键值对

上一篇:remote指令添加远程数据库


下一篇:currentColor-CSS3非常有用的变量