亚洲精品无码专区2,亚洲处破女AV日韩精品,国产成人精品免费视频动漫

首字母大小寫無關模式
有一段時間，我在寫正則表達式來匹配Drug關鍵字時，經常寫出 /viagra|cialis|anti-ed/ 這樣的表達式。為了讓它更美觀，我會給關鍵詞排序；為了提升速度，我會使用 /[Vv]iagra/ 而非/viagra/i ，只讓必要的部分進行大小寫通配模式。確切地說，我是需要對每個單詞的首字母進行大小寫無關的匹配。

我寫了這樣的一個函數，專門用來批量轉換。

復制代碼代碼如下:

#convert regex to sorted list, then provide both lower/upper case for the first letter of each word
#luf means lower upper first

sub luf{
# split the regex with the delimiter |
my @arr=sort(split(/|/,shift));

# provide both the upper and lower case for the
# first leffer of each word
foreach (@arr){s/b([a-zA-Z])/[l$1u$1]/g;}

# join the keyword to a regex again
join(‘|’,@arr);
}

控制全局匹配下次開始的位置

記得jyf曾經問過我，如何控制匹配開始的位置。嗯，現在我可以回答這個問題了。Perl 提供了 pos 函數，可以在 /g 全局匹配中調整下次匹配開始的位置。舉例如下：

復制代碼代碼如下:

$_=”abcdefg”;
while(/../g)
{
print $&;
}

其輸出結果是每兩個字母，即ab, cd, ef

可以使用 pos($_)來重新定位下一次匹配開始的位置，如：

復制代碼代碼如下:

$_=”abcdefg”;
while(/../g)
{
pos($_)–; #pos($_)++;
print $&;
}

輸出結果：

復制代碼代碼如下:

pos($_)–: ab, bc, cd, de, ef, fg.
pos($_)++: ab, de.

可以閱讀 Perl 文檔中關于 pos的章節獲取詳細信息。

散列與正則表達式替換
《effective-perl-2e》第三章有這樣一個例子（見下面的代碼），將特殊符號轉義。

復制代碼代碼如下:

my %ent = { ‘&’ => ‘amp’, ‘<‘ => ‘lt’, ‘>’ => ‘gt’ };
$html =~ s/([&<>])/&$ent{$1};/g;

這個例子非常非常巧妙。它靈活地運用了散列這種數據結構，將待替換的部分作為 key ，將與其對應的替換內容作為 value 。這樣只要有匹配就會捕獲，然后將捕獲的部分作為 key ，反查到 value 并運用到替換中，體現了高級語言的效率。

不過，這樣的 Perl 代碼，能否移植到 Python 中呢？ Python 同樣支持正則，支持散列（Python 中叫做 Dictionary），但是似乎不支持在替換過程中插入太多花哨的東西（替換行內變量內插）。

查閱 Python 的文檔，（在 shell 下執行 python ，然后 import re，然后 help(re)），：

復制代碼代碼如下:

sub(pattern, repl, string, count=0)
Return the string obtained by replacing the leftmost
non-overlapping occurrences of the pattern in string by the
replacement repl. repl can be either a string or a callable;
if a string, backslash escapes in it are processed. If it is
a callable, it’s passed the match object and must return
a replacement string to be used.

原來 python 和 php 一樣，是支持在替換的過程中使用 callable 回調函數的。該函數的默認參數是一個匹配對象變量。這樣一來，問題就簡單了：

復制代碼代碼如下:

ent={‘<‘:”lt”,
‘>’:”gt”,
‘&’:”amp”,
}

def rep(mo):
return ent[mo.group(1)]

html=re.sub(r”([&<>])”,rep, html)

python 替換函數 callback 的關鍵點在于其參數是一個匹配對象變量。只要明白了這一點，查一下手冊，看看該種對象都有哪些屬性，一一拿來使用，就能寫出靈活高效的 python 正則替換代碼。

一	二	三	四	五	六	日
« 6月
	1	2	3	4	5	6
7	8	9	10	11	12	13
14	15	16	17	18	19	20
21	22	23	24	25	26	27
28	29	30	31

正則表達式筆記三則

相關推薦

熱門標簽

近期文章