Go to file

Ilya Kreymer 2717a60763 improvements / bug fixes for stop/cancel handling: (#279 ) - only send signal if stopping, no need for canceling as pods/containers will be removed - refactor stop/cancel handling to be unified in manager, separate in job - when stopping / graceful shutdown, return false if sending signal fails - return success=true in json response if and only if stop/cancel actually succeeds, return 'error' message in error, should fix #270 - allow canceling after stopping / if stopping fails - ensure finished time is set in case of cancelation before crawl starts, should fix #273		2022-06-29 17:47:25 -07:00
.github/workflows	Local swarm + podman support (#261 )	2022-06-14 00:13:49 -07:00
backend	improvements / bug fixes for stop/cancel handling: (#279 )	2022-06-29 17:47:25 -07:00
chart	nginx simplify: (#259 )	2022-06-13 11:53:15 -07:00
configs	config sample: switch back to browsertrix-crawler:latest for now	2022-06-17 13:39:45 -07:00
frontend	improvements / bug fixes for stop/cancel handling: (#279 )	2022-06-29 17:47:25 -07:00
scripts	config/scripts:	2022-06-16 22:36:44 -07:00
test	Single config and env vars (#267 )	2022-06-16 21:50:03 -07:00
.gitignore	Local swarm + podman support (#261 )	2022-06-14 00:13:49 -07:00
Deployment.md	Single config and env vars (#267 )	2022-06-16 21:50:03 -07:00
docker-compose.yml	Single config and env vars (#267 )	2022-06-16 21:50:03 -07:00
LICENSE	Add License, Logo and README updates for release (#157 )	2022-02-23 12:10:46 -08:00
NOTICE	Add License, Logo and README updates for release (#157 )	2022-02-23 12:10:46 -08:00
pylintrc	misc tweaks:	2021-08-25 18:34:49 -07:00
README.md	Local swarm + podman support (#261 )	2022-06-14 00:13:49 -07:00

README.md

Browsertrix Cloud

Browsertrix Cloud is an open-source cloud-native high-fidelity browser-based crawling service designed to make web archiving easier and more accessible for everyone.

The service provides an API and UI for scheduling crawls and viewing results, and managing all aspects of crawling process. This system provides the orchestration and management around crawling, while the actual crawling is performed using Browsertrix Crawler containers, which are launched for each crawl.

The system is designed to run in both Kubernetes and Docker Swarm, as well as locally under Podman.

See Features for a high-level list of planned features.

Deployment

See the Deployment page for information on how to deploy Browsertrix Cloud.

Development Status

Browsertrix Cloud is currently in an alpha stage and not ready for production. This is an ambitious project and there's a lot to be done!

If you would like to help in a particular way, please open an issue or reach out to us in other ways.

License

Browsertrix Cloud is made available under the AGPLv3 License.

If you would like to use it under a different license or have a question, please reach out as that may be a possibility.