If you can’t boat, dream…
As is almost everyone in the world right now, we are stuck and are having difficulty even imaging getting out on the boat in the foreseeable future. So to amuse myself I have been looking online for that future “forever” boat.
Who am I kidding; I am always looking online at boats. But what I have done is work on a project that fixes an annoyance of mine.
Some time in the recent past YachtWorld, which is the de facto standard for listing boats, decided to redo their website. And one of the outcomes of that is that you can no longer search for boats in multiple places at the same time. And seeing how I desire to buy a boat in the PNW (which used to be a selectable search category) and the PNW ostensibly includes British Columbia, Washington state and Oregon, I now had to perform three separate searches with no way to “save” a previous search and be able to compare. A definite downgrade if you ask me.
I had recently started to learn how to scrape websites using python because I wanted to be able to post a “books read” page on my personal site. Web scraping is a method of visiting a site and stripping the data from it. It occurred to me that I could apply the same procedure to YachtWorld and retrieve and aggregate all three searches to one output. So off I went.
Building
My initial method was to build a python script that output the data to a php web page. After a lot of trial and error it worked. But this meant every time I wanted to do a new search I would have to run the terminal command again and enter in the price range (which were the only variables I included). This produced a database file which the php page could then display. This was pretty onerous, particularly because I kept forgetting the proper commands :-). I also held out some hope that eventually I could share this with other people. There had to be a better way.
So I investigated further and discovered Flask, which is—simplistically put—a way to turn a python script into a web app. After a lot of fiddling (which I will detail soon over at macblaze.ca) I came up with this:
The new version added the options for Canadian vs. U.S. currency and length and the ability to sort the results. It is all one compact, self-contained unit and works 100% from a web browser. The search returns most of the details of the boats and links back to the official YachtWorld result so you can dive deeper into any particular boat. If you ask me it turned out pretty good.
A qualified success.
Deployment
I am stuck now though. It currently runs on my testing server and when I went looking to deploy it, discovered that my official web host doesn’t provide python support. I would either have to switch providers or get an additional (paid) account at another web host. I did find pythonanywhere.com, which offered free, albeit limited, python/flask hosting and was excited for the full hour it took to set up and get it running. You can find it here (http://btk.pythonanywhere.com/) but unfortunately the free accounts won’t let you scrape websites that don’t have official api’s (application programming interface)—which YachtWorld does, but it isn’t free so I don’t have access to it. So while the app is running it won’t actually deliver any results.
So now I am faced with the choice to either open up my testing server to the world (which I am not likely to do), change providers (and I just paid for the next couple of years) or sit on my little project and declare it for personal use only. I haven’t given up yet, but the prospects look dim.
So…
So here I sit, a vast distance from the boat and any prospect of cruising and a weeks worth of work sitting on a computer with no way to share. And I am now getting tired of running the search over and over again. Sigh.
Anyway, if any of you experience the same frustration with YachtWorld, I sympathize and want you to know that there is a way around it with a bit of time and even less knowledge—which is something we should all have in abundance right now.
—Bruce #Purchasing
Instagram This Week
Instagram This Week
Instagram This Week
What kind of cruiser?
Bored of the lack of thoughtfulness and reading comprehension on a cruising forum I frequent I post the following post (the original post). I got some good feedback on my attempt at humour so I thought I would repost it and a few of my follow-up posts…
Warning: the following is mostly tongue-in-cheek, although there is an underlying truth to the issue.
It seems to me that many of the discussions on CF go off the rails because we are all so different in both our ambitions and realities. “How much does it cost to cruise?” You’d think this was a pretty straightforward question but it really isn’t and the topic generally goes sideways faster than a drunken docking attempt on a windy day. Even seemingly innocuous discussions like “How much chain do I need in such-and-such area” often become completely nonsensical because no one can agree on what we are actually talking about.
It seems to me we all stubbornly live in our own bubbles. Of course he means to be on the hook 99% of the time! vs. No sensible person would ever anchor out if there was a marina nearby! is never considered when joining the discussion—we just dive in and start the pontificating.
So I propose we find a solution.
I started with a simple chart. “That,” I said to myself, “would solve everything.” “It’s simple Self,” I said, “Simply find your place on the chart and colour code your answers so as to clear up any potential misunderstandings with all our fellow CFers.”

But then I got to thinking. It just wasn’t enough. Where was the geography? We needed a plane for warm vs cold, anchorages vs moorings… And what about round-the-world vs regional? Or a scale for single-handers all the way up to large crew?
(I tried to make a chart—it turned into a disaster )

This was getting complicated. So maybe a chart is out. How about codes?
Crew size |Crew Period |Climate | Expenses | Area | Moorage Style etc.
So for example I would be a 2PTCMRA (2 crewed, part-time, cold-water, moderate spending, regional anchorer). We could have a big chart … or an app…this is a great opportunity for an app developer—wait…ummm, we need a code for electronics too, with sub codes for radios, GPS, trackers, weather systems… hmmm…
So… ya. A chart. A BIG chart! We will need codes for anchor types too, and gun preferences, probably sub charts for specific regions cause we all know that the Med is different kettle of fish than the Caribbean— I mean literally, totally different fish and what we have here in the PNW is obviously so much better than anywhere else so…but I digress.
We will need indicators for partiers, outgoing folk, introverts, singles, people who can rebuild transmissions blindfolded, those who are still unsure what the pointy bits on a hammer are for… Oh, and most especially some way to differentiate where we fall on the scale of way over-prepared to “I bought a boat on the internet and cast off tomorrow—can someone show me how to sail?”
I guess we need a code for sailing purists. And one for those of us who have a funny stick in the middle of our powerboats. And racers. And dock-bound liveaboards. Is specifying whale preferences going too far? I like orca myself. What the hell let’s add it in.
So now would start any thread participation with “I am a 2PTCMRAPNWNGIMIEMPSO…” and as a result any potential misunderstandings would immediately be averted.
Ok so that about covers it. I am starting to grid this out and will announce the official chart when its completed and will have to get the mods onboard to enforce total compliance. Won’t work otherwise and this would have been a total waste of effort. Hmmm, I guess we will have to access penalties as well. Maybe a grid to determine just how serious breaches, cases of mis-information, and what the levels of incompleteness are and empower Guardians of the Grid to deliver appropriate punishments. That can be phase 2.
This is going to be so cool!
Or… I guess… We could…I mean… Maybe…
Could we all just take a moment to thank god (and/or whatever diety, spirit or scientific principle you believe in) that we are all able to get out on the water in whatever way we can and that that is a whole lot of different ways?
Seriously, if we all just took a moment to consider an OP’s situation or even a fellow thread participant’s perspective there would be so many fewer stupid, inadvertent pissing contests. And then we could get on with the intentional ones.
Just a thought. Let me know if you still want me to go ahead with the grid idea.
The thread continued. Both Mike O’ and L independently thought Venn diagrams were the way to go:
Maybe what you need is a series of Venn diagrams. Cruisers are liveaboards, but not all liveaboards are cruisers. Racers and cruisers overlap, but some are just racers, and some are just cruisers. Some cruisers anchor out, some marina hop, and some do both.
So I gave this a try:

Then Wolfgal suggested:
What fun, Macblaze! Your ever-growing charts are so complicatedly fun! a holograph chart that actually turns around in space, popping up from my personal R2-D2 would take your idea to the next level. could you do this?. 
I thought, I can do that…

I almost lost control of the whole idea when Tayana42 tried to outsmart me with:
Curiosity
Reigns
Until
I’ve
Seen
Enough
Or
Casually
Roving
Upon
Iceless
Seas
In
No
Great hurry
Clever, clever! It went on for pages and pages after that.
Instagram This Week
Instagram This Week
Instagram This Week
Instagram This Week
Web scraping Python code
In my previous post I explained that I was looking for a way to use web scraping to extract data from my Calibre-Web shelves and automatically post them to my Books Read page here on my site. In this post I will step through my final Python script to explain to my future self what I did and why.
Warning: Security
A heads up. I have no guarantees that this code is secure enough to use in a production environment. In fact I would guess it isn’t. But my Web-Calibre webserver is local to my home network and I trust that my hosted server (macblaze.ca) is secure enough. But since you are passing passwords etc. back and forth I wouldn’t count on any of this to be secure without a lot more effort than I am willing to put in.
The code in bits
# import various libraries
import requests
from bs4 import BeautifulSoup
import re
This loads the various libraries the script uses. Requests is a http library that allows you to send requests to websites, BeautifulSoup is a library to pull data from html and re is a regex library to allow you to do custom searches.
# set variables
# set header to avoid being labeled a bot
headers = {
'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36'
}
# set base url
urlpath='http://urlpath'
# website login data
login_data = {
'next': '/',
'username': 'username',
'password': 'password',
'remember_me': 'on',
}
# set path to export as markdown file
path_folder="/Volumes/www/books/"
file = open(path_folder+"filename.md","w")
This sets up various variables used for login including a header to try and avoid being labeled a bot, the base url of the Calibre-web installation, login data and specifies a location and name for the resulting markdown file. The open command is marked with a ‘w’ switch to indicate the script will write a new file every time it is executed, overwriting the old one.
# log in and open http session
with requests.Session() as sess:
url = urlpath+'/login'
res = sess.get(url, headers=headers)
res = sess.post(url, data=login_data)
Then, using Requests, I opened a session on the webserver and log in using the variables.
Writing the File
Note: The code has matching file.write() and print() statements throughout. The print() statements just write to the terminal app and allow me to see what is being written to the actual file using file.write(). They are completely unnecessary.
# Set Title
file.write("# Books Read\n")
print("# Books Read\n")
Pretty basic: write the words Books Read followed by a carriage return, tagged with a # to indicate it is a h1 head. This will become the actual page name.
# find list of shelves
shelfhtml = sess.get(urlpath)
soup = BeautifulSoup(shelfhtml.text, "html.parser")
shelflist = soup.find_all('a', href=re.compile('/shelf/[1-9]'))
print (shelflist)
So now we set the variable shelfhtml to the session we opened earlier. Using BeautifulSoup we grab all the html code and search for all a links that have an href that contain the regex expression ‘/shelf/[1-9]’. (Hopefully I won’t have more than 9 shelves or I will have to redo this bit.) The variable now contains list of all the links that match that pattern and looks like this:
[<a href="/shelf/3"><span class="glyphicon glyphicon-list private_shelf"></span>2018</a>, <a href="/shelf/2"><span class="glyphicon glyphicon-list private_shelf"></span>2019</a>, <a href="/shelf/1"><span class="glyphicon glyphicon-list private_shelf"></span>2020</a>]
This as you can see, contains the links to all three of my current Year shelves, displayed in ascending numerical order.
#reverse order of urllist
dateshelflist=(get_newshelflist())
dateshelflist.reverse()
print (dateshelflist)
I wanted to display my book lists from newest to oldest so I used python to reverse the items in the list.
First loop: the shelves
The first loop loops through all the shelves (in this case 3 of them) and starts the process of building a book list for each.
# loop through sorted shelves
for shelf in dateshelflist:
#set shelf page url
res = sess.get(urlpath+shelf.get('href'))
soup = BeautifulSoup(res.text, "html.parser")
# find year from shelflist and format
shelfyear = soup.find('h2')
year = re.search("([0-9]{4})", shelfyear.text)
year.group()
file.write("### {}\n".format(year.group()))
print("### {}\n".format(year.group()))
In the first iteration of the loop, the script goes to the actual shelf page using the base url and then adding an href extracted from the list by using a get command and then accesses the html from the resulting webpage. Then the script finds the year info, which is a H2, extracts the 4-digit year with the regex ([0-9]{4}) and writes it to the file, formatted as an H3 header and followed by a line break.
# find all books
books = soup.find_all('div', class_='col-sm-3 col-lg-2 col-xs-6 book')
Using BeautifulSoup we extract the list of books from the page knowing they are all marked with a div in the class col-sm-3 col-lg-2 col-xs-6 book.
Second loop: the books
#loop though books. Each book is a new BeautifulSoup object.
for book in books:
title = book.find('p', class_='title')
author = book.find('a', class_='author-name')
seriesname = book.find('p', class_='series')
pubdate = book.find('p', class_='publishing-date')
coverlink = book.find('div', class_='cover')
if None in (title, author, coverlink, seriesname, pubdate):
continue
# extract year from pubdate
pubyear = re.search("([0-9]{4})", pubdate.text)
pubyear.group()
This is the beginning of the second loop. For each book we use soup to extract the title, author, series, pubdate and cover (which I don’t end up using). Each search is based on the class assigned to it in the original html code. Because I only want the pub year and not pub date, I again use a regex to extract the 4-digit year. The if None… statement is there just in case one of the fields is empty and prevents the script from hanging.
# construct line using markdown
newstring = "* ***{}*** — {} ({})\{} – ebook\n".format(title.text, author.text, pubyear.group(), seriesname.text)
file.write(newstring)
print (newstring)
Next we construct the book entry based on how we want it to appear on the web page. In my case I want each entry to be an li and end up looking like this:
- The Cloud Roads — Martha Wells (2011)
Book 1.0 of Raksura – ebook
Python allows you to just list the variables at the end of the statement and fills in the {} automatically which makes for easier formatting. The script then writes the line to the open markdown file and heads up to the beginning of the loop to grab the next book.
More loops
That’s pretty much it. It loops through the books until it runs out and heads back to the first loop to see if there is another shelf to process. After it processes all the shelves it drops to the last line of the script:
file.close()
which closes the file and that is that—c’est tout. It will now be accessed the next time some visits the Books Read page on my site.
In Conclusion
Hopefully this is clear enough so that when I forget every scarp of python in the years to come I can still recreate this after the inevitable big crash. The script, called scrape.py in my case, is executed in terminal by going to the enclosing folder and typing python3 scrape.py then hitting enter. Automating that is something I will ponder if this book list thing becomes my ultimate methodology for recording books read. It’s big failing is that it only records ebooks in my Calibre library. I might have to redo the entire thing for something like LibraryThing where I can record all my books…lol. Hmmm… maybe…
The Final Code
Here is the final script in its entirety.
# import various libraries
import requests
from bs4 import BeautifulSoup
import re
# set header to avoid being labeled a bot
headers = {
'user-agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36'
}
# set base url
urlpath='http://urlpath'
# website login data
login_data = {
'next': '/',
'username': 'username',
'password': 'password',
'remember_me': 'on',
}
# set path to export as markdown file
path_folder="/Volumes/www/home/books/"
file = open(path_folder+"filename.md","w")
with requests.Session() as sess:
url = urlpath+'/login'
res = sess.get(url, headers=headers)
res = sess.post(url, data=login_data)
# Note: print() commands are purely for terminal output and unnecessary
# Set Title
file.write("# Books Read\n")
print("# Books Read\n")
# find list of shelves
shelfhtml = sess.get(urlpath)
soup = BeautifulSoup(shelfhtml.text, "html.parser")
shelflist = soup.find_all('a', href=re.compile('/shelf/[1-9]'))
# print (shelflist)
#reverse order of urllist
dateshelflist=(get_newshelflist())
dateshelflist.reverse()
# print (dateshelflist)
# loop through sorted shelves
for shelf in dateshelflist:
#set shelf page url
res = sess.get(urlpath+shelf.get('href'))
soup = BeautifulSoup(res.text, "html.parser")
# find year and format
shelfyear = soup.find('h2')
year = re.search("([0-9]{4})", shelfyear.text)
year.group()
file.write("### {}\n".format(year.group()))
print("### {}\n".format(year.group()))
# find all books
books = soup.find_all('div', class_='col-sm-3 col-lg-2 col-xs-6 book')
#loop though books. Each book is a new BeautifulSoup object.
for book in books:
title = book.find('p', class_='title')
author = book.find('a', class_='author-name')
seriesname = book.find('p', class_='series')
pubdate = book.find('p', class_='publishing-date')
coverlink = book.find('div', class_='cover')
if None in (title, author, coverlink, seriesname, pubdate):
continue
# extract year from pubdate
pubyear = re.search("([0-9]{4})", pubdate.text)
pubyear.group()
# construct line using markdown
newstring = "* ***{}*** — {} ({})\{} – ebook\n".format(title.text, author.text, pubyear.group(), seriesname.text)
file.write(newstring)
print (newstring)
file.close()
Note 12/2021
There has been an update to the Calibre web code so I had to make some changes to the python script.











