python 匹配中文和英文

时间：2015-07-21 20:31:47 阅读：194 评论：0 收藏：0 [点我收藏+]

在处理文本时经常会匹配中文名或者英文word，python中可以在utf-8编码下方便的进行处理。

中文unicode编码范围[\u4e00-\u9fa5]

英文字符编码范围[a-zA-Z]

此时匹配连续的中文或者英文就很方便了，例如：

>>> import re
>>> strings = u‘中国china美国American‘
>>> print strings
中国china美国American
>>> ch_pat = re.compile(ur‘[\u4e00-\u9fa5]+‘)
>>> en_pat = re.compile(‘[a-zA-Z]+‘)
>>> ch_words = ch_pat.findall(strings)
>>> en_words = en_pat.findall(strings)
>>> print ch_words
[u‘\u4e2d\u56fd‘, u‘\u7f8e\u56fd‘]
>>> print en_words
[u‘china‘, u‘American‘]

python 匹配中文和英文

原文：http://www.cnblogs.com/chybot/p/4665389.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年09月23日 (328)
2021年09月24日 (313)
2021年09月17日 (191)
2021年09月15日 (369)
2021年09月16日 (411)
2021年09月13日 (439)
2021年09月11日 (398)
2021年09月12日 (393)
2021年09月10日 (160)
2021年09月08日 (222)