1236. Web Crawler
Problem
Starting from a URL, crawl all URLs reachable from it that share the same hostname.
Given a starting URL startUrl, perform a web crawl. Starting from startUrl, crawl all URLs that can be visited by following at most one hop from the current page. A URL follows this format: http://{hostname}/path. Two URLs belong to the same hostname if the part after "http://" and before the first "/" are identical.
Return all URLs (including startUrl) that share the same hostname with startUrl.
Examples
Input: ["http://news.google.com/about",["http://news.google.com","http://news.google.com/videos"]]
Output: ["http://news.google.com","http://news.google.com/about","http://news.google.com/videos"]
Input: ["http://a.com/a",["http://a.com/a","http://a.com/b"]]
Output: ["http://a.com/a","http://a.com/b"]
Hints
Extract the hostname from startUrl using string parsing.
Use BFS or DFS to explore URLs.
Only add URLs to the crawl list if they share the same hostname.
Related Problems
1236. Web Crawler
Starting from a URL, crawl all URLs reachable from it that share the same hostname.