ES查询索引字段的分词结果

这篇具有很好参考价值的文章主要介绍了ES查询索引字段的分词结果。希望对大家有所帮助。如果存在错误或未考虑完全的地方,请大家不吝赐教,您也可以点击"举报违法"按钮提交疑问。

一、_termvectors 

1、查看文档中某一个字段的分词结果

GET /{index}/{type}/{_id}/_termvectors?fields=[field]

2、样例:

text的值为:https://www.b4d99.com/html/202204/45672.html

GET http://IP:POST/textcontent_2022/textcontent/20220422191235893045256250/_termvectors?fields=text

得到的结果:

"terms": {
	"202204": {
		"term_freq": 1,
		"tokens": [
			{
				"position": 4,
				"start_offset": 27,
				"end_offset": 33
			}
		]
	},
	"45672": {
		"term_freq": 1,
		"tokens": [
			{
				"position": 5,
				"start_offset": 34,
				"end_offset": 39
			}
		]
	},
	"com": {
		"term_freq": 1,
		"tokens": [
			{
				"position": 2,
				"start_offset": 18,
				"end_offset": 21
			}
		]
	},
	"html": {
		"term_freq": 2,
		"tokens": [
			{
				"position": 3,
				"start_offset": 22,
				"end_offset": 26
			},
			{
				"position": 6,
				"start_offset": 40,
				"end_offset": 44
			}
		]
	},
	"https": {
		"term_freq": 1,
		"tokens": [
			{
				"position": 0,
				"start_offset": 0,
				"end_offset": 5
			}
		]
	},
	"www.b4d99": {
		"term_freq": 1,
		"tokens": [
			{
				"position": 1,
				"start_offset": 8,
				"end_offset": 17
			}
		]
	}
}

二、_analyze

1、语法

POST _analyze
{
  "analyzer": "具体的分词器",
  "text": "待分词的内容"
}

2、样例:

text的值为:https://www.b4d99.com/html/202204/45672.html

POST _analyze
{
  "analyzer": "standard",
  "text": "https://www.b4d99.com/html/202204/45672.html"
}

得到的结果:文章来源地址https://www.toymoban.com/news/detail-504858.html

{
    "tokens": [
        {
            "token": "https",
            "start_offset": 0,
            "end_offset": 5,
            "type": "<ALPHANUM>",
            "position": 0
        },
        {
            "token": "www.b4d99",
            "start_offset": 8,
            "end_offset": 17,
            "type": "<ALPHANUM>",
            "position": 1
        },
        {
            "token": "com",
            "start_offset": 18,
            "end_offset": 21,
            "type": "<ALPHANUM>",
            "position": 2
        },
        {
            "token": "html",
            "start_offset": 22,
            "end_offset": 26,
            "type": "<ALPHANUM>",
            "position": 3
        },
        {
            "token": "202204",
            "start_offset": 27,
            "end_offset": 33,
            "type": "<NUM>",
            "position": 4
        },
        {
            "token": "45672",
            "start_offset": 34,
            "end_offset": 39,
            "type": "<NUM>",
            "position": 5
        },
        {
            "token": "html",
            "start_offset": 40,
            "end_offset": 44,
            "type": "<ALPHANUM>",
            "position": 6
        }
    ]
}

到了这里,关于ES查询索引字段的分词结果的文章就介绍完了。如果您还想了解更多内容,请在右上角搜索TOY模板网以前的文章或继续浏览下面的相关文章,希望大家以后多多支持TOY模板网!

本文来自互联网用户投稿,该文观点仅代表作者本人,不代表本站立场。本站仅提供信息存储空间服务,不拥有所有权,不承担相关法律责任。如若转载,请注明出处: 如若内容造成侵权/违法违规/事实不符,请点击违法举报进行投诉反馈,一经查实,立即删除!

领支付宝红包 赞助服务器费用

相关文章

觉得文章有用就打赏一下文章作者

支付宝扫一扫打赏

博客赞助

微信扫一扫打赏

请作者喝杯咖啡吧~博客赞助

支付宝扫一扫领取红包,优惠每天领

二维码1

领取红包

二维码2

领红包