JavaScript removing HTML tags

I'm a full-stack developer from South Africa 🇿🇦. I love writing about JavaScript, HTML and CSS.
Search for a command to run...

I'm a full-stack developer from South Africa 🇿🇦. I love writing about JavaScript, HTML and CSS.
No comments yet. Be the first to comment.
Most of you know me for my consistency, a golden arrow in my blog series. I've written 1000 articles in 1008 days! Almost an article a day, and my honeymoon was the only holiday I ever took. I'm super proud of this achievement; it has been a fantasti...

It's not the first time I'll be talking about community. I think it's an essential aspect of any successful tool. This shows in my previous explorations of Astro, Medusa, and now Vendure as well. All these products thrive in a super open, welcoming, ...

The cool part about Vendure is how easy it is to set up and how abstract each layer is. Basically, we get the following elements: External database Server Worker Admin UI Frontend While this is amazing, it also brings a bit of complexity when it co...

The previous article looked at customizing Vendure on a data and process level. In this article, we'll look at customizing emails, as they are often a big part of a webshop system. We'll be looking at two different layers of customization for customi...

Even though Vendure is a pretty significant project out of the box, in some cases, we might want to go in and modify some elements to work to our specific use case. In this article, I'll take a high-level look at some elements we can customize within...

I recently needed to remove all HTML from the content of my own application.
In this case, it was to share a plain text version for meta descriptions, but it can be used for several outputs.
Today I'll show you two ways of doing this, which are not fully safe if your application accepts user inputs.
Users love to break scripts like this and especially method one can give you some vulnerabilities.
One method is to create a temporary HTML element and get the innerText from it.
const original = `<h1>Welcome to my blog</h1>
<p>Some more content here</p><br /><img alt="a > 2" src="img.jpg" />`;
let removeHTML = input => {
let tmp = document.createElement('div');
tmp.innerHTML = input;
return tmp.textContent || tmp.innerText || '';
}
console.log(removeHTML(original));
This will result in the following:
'Welcome to my blog
Some more content here'
As you can see we removed every HTML tag including a bogus image.
My personal favourite for my own applications is using a regex, just a cleaner solution and I trust my own inputs to be valid HTML.
How it works:
const original = `<h1>Welcome to my blog</h1>
<p>Some more content here</p><br /><img src="img.jpg" />`;
const regex = original.replace(/<[^>]*>/g, '');
console.log(regex);
This will result in:
'Welcome to my blog
Some more content here'
As you can see, we removed the heading, paragraph, break and image.
This is because we escape all < > formats.
It could be breached by something silly like:
const original = `<h1>Welcome to my blog</h1>
<p>Some more content here</p><br /><img alt="a > 2" src="img.jpg" />`;
I know it's not valid HTML anyhow and one should use > for this.
But running this will result in:
'Welcome to my blog
Some more content here 2" src="img.jpg" />'
It's just something to be aware of.
You can have a play with both methods in this Codepen.
Thank you for reading my blog. Feel free to subscribe to my email newsletter and connect on Facebook or Twitter