我正在嘗試從文字檔案中的 URL 清單中替換 HTML 檔案中的所有圖像來源 URL。
文件1.html
<td class="MetadataRes" width="380px" colspan="2" style="border-top: 1px #336699 solid;">
<a olv_link="/Default/Scripting/ArticleWin.asp?From=Search&Key=Orange/2011/03/27/129/Ad12911.xml&CollName=Orange_APA3&DOCID=2485870&PageLabelPrint=H2&Skin=%4f%72%61%6e%67%65%43%6f%75%6e%74%79%52%65%67%69%73%74%65%72&AW=%31%34%31%32%36%32%38%32%31%34%35%30%32&sPublication=%4f%72%61%6e%67%65&sScopeID=%44%52&SECTION=%43%6c%61%73%73%69%66%69%65%64&sSorting=%53%63%6f%72%65%2c%64%65%73%63&sQuery=%72%65%67%69%73%74%65%72%65%64%20%6e%75%72%73%65%20%3c%4f%52%3e%20%52%4e&rEntityType=&sSearchInAll=%66%61%6c%73%65&sDateFrom=%25%33%30%25%33%35%25%32%66%25%33%30%25%33%31%25%32%66%25%33%32%25%33%30%25%33%31%25%33%30&sDateTo=%25%33%30%25%33%35%25%32%66%25%33%33%25%33%31%25%32%66%25%33%32%25%33%30%25%33%31%25%33%31&dc:creator=&PageLabel=&dc:publisher=&RefineQueryView=&StartFrom=%30" href="javascript:void(0);" onclick="window.top.sys.openArtWin(this.getAttribute('Olv_link'))">
<img src="/Repository/GetImage.dll?baseHref=Orange/2011/03/27&EntityID=Ad12911&imgExtension=">
</a>
</td>...
* 請在此處查看完整文件:http://pastebin.com/XbwtZJPa
文件2.txt
/getimage.dll?path=Orange/2011/03/27/129/Img/Ad1291103.gif
/getimage.dll?path=Orange/2011/03/20/133/Img/Ad1330402.gif
/getimage.dll?path=Orange/2010/08/29/137/Img/Ad1372408.gif
我想將上述 HTML 文件中圖像的 URL 替換為 URL 文件中列出的第一個 URL,以獲得以下內容:
結果.html
<td class="MetadataRes" width="380px" colspan="2" style="border-top: 1px #336699 solid;">
<a olv_link="/Default/Scripting/ArticleWin.asp?From=Search&Key=Orange/2011/03/27/129/Ad12911.xml&CollName=Orange_APA3&DOCID=2485870&PageLabelPrint=H2&Skin=%4f%72%61%6e%67%65%43%6f%75%6e%74%79%52%65%67%69%73%74%65%72&AW=%31%34%31%32%36%32%38%32%31%34%35%30%32&sPublication=%4f%72%61%6e%67%65&sScopeID=%44%52&SECTION=%43%6c%61%73%73%69%66%69%65%64&sSorting=%53%63%6f%72%65%2c%64%65%73%63&sQuery=%72%65%67%69%73%74%65%72%65%64%20%6e%75%72%73%65%20%3c%4f%52%3e%20%52%4e&rEntityType=&sSearchInAll=%66%61%6c%73%65&sDateFrom=%25%33%30%25%33%35%25%32%66%25%33%30%25%33%31%25%32%66%25%33%32%25%33%30%25%33%31%25%33%30&sDateTo=%25%33%30%25%33%35%25%32%66%25%33%33%25%33%31%25%32%66%25%33%32%25%33%30%25%33%31%25%33%31&dc:creator=&PageLabel=&dc:publisher=&RefineQueryView=&StartFrom=%30" href="javascript:void(0);" onclick="window.top.sys.openArtWin(this.getAttribute('Olv_link'))">
<img src="/Repository/getimage.dll?path=Orange/2011/03/27/129/Img/Ad1291103.gif">
</a>
</td>...
有推薦的 shell 命令來執行此操作嗎?我在運行 10.9 的 Mac 上考慮了以下 sed 命令,但遇到了錯誤。
$ gsed -e 's/.*SRC="\/Repository\([^"]*\)".*/\1/p{r File1.html' -e 'd}' File2.txt
答案1
假設包含EntityID
一個唯一的字串來識別 File2.txt 中的正確 URL,這不僅適用於您的範例:
sed '\_^/getimage.dll.*gif$_{H;d}
G;s/<img src="[^"]*EntityID=\([^&]*\)&[^"]*"\(.*\)\n\(\/getimage[^\n]*\1[^.]*.gif\).*/<img src="\3"/;s/\n.*//' File2.txt File1.html
如果需要,請尋求解釋。