Let's say I have the following HTML files:
html1.html
<html>
<head>
<link href="blah.css" rel="stylesheet" type="text/css" />
</head>
<body>
<div>this here be a div, y'all</div>
</body>
</html>
html2.html
<html>
<head>
<script src="blah.js"></script>
</head>
<body>
<span>this here be a span, y'all</span>
</body>
</html>
I want to take these two files and make a master file that would look like this:
<html>
<head>
<link href="blah.css" rel="stylesheet" type="text/css" />
<script src="blah.js"></script>
</head>
<body>
<div>this here be a div, y'all</div>
<span>this here be a span, y'all</span>
</body>
</html>
Is this possible using a simple Linux command? I've tried looking at join, but it looks like that joins on a common field, and I'm not necessarily going to have common fields... I just need to basically add the difference, but also have the main structure still intact (I guess this could be referred to as a left-join?). Doesn't look like cat will work either... as that merges by appending one file, then the next, etc.
If there isn't a simple Linux command, my next step is to either write a script that compares both scripts line by line, or create a master HTML file that references these two individual files somehow.
5 Answers
Use pandoc to merge e.g. all html-files in the current directory:
pandoc -s *.html -o output.html
You can use html-merge tool to merge multiple HTML files preserving their internal hypertext links. It's a win32 program, but you can run it in linux using Wine. Download page:
Your example files are well-formed XHTML. Excellent! This means you can use a simple XSLT script. See How to merge two XML files with XSLT
Here is a simple solution that uses Python's lxml library, though it will only copy element children of the body tag selected child::*, not text nodes, which would require a modification child::node() and some extra logic for dealing with appending text nodes.
#!/usr/bin/python3
import sys, os
from lxml.html import tostring, parse
if len(sys.argv) < 2:
print("Usage: merge.py [file1] ... [filen] [outfile]")
if os.path.isfile(sys.argv[-1]):
if input('Override? (y/n) ' + sys.argv[-1]) != 'y':
sys.exit(0)
def tostr(n):
try:
return tostring(n)
except:
return str(n)
tree = parse(sys.argv[1])
for f in sys.argv[2:-1]:
print(f)
tree2 = parse(f)
for n in tree2.xpath('//head/child::*'):
if all([tostr(n) != tostr(n2)\
for n2 in tree2.xpath('//head/child::*')]):
tree.xpath('//head')[0].append(n)
for n in tree2.xpath('//body/child::*'):
tree.xpath('//body')[0].append(n)
tree.write(sys.argv[-1])
Save this to a file merge.py and run chmod +x merge.py.
Usage: merge.py [file1] ... [filen] [outfile]
If it fails, one or more files are malformed and need to be fixed either manually or with htmllint or hxnormalize.
The quickest way I found with out using any other programs is:
cat html2.html >> html1.html
This will add html2.html to the end of html1.html or if you want both of them in a new file you can type
cat html1.html >> html3.html && cat html2.html >> html3.html
Into the terminal. The >> appends the code in the file to the other code in the other one.