browsertrix

Author	SHA1	Message	Date
Tessa Walsh	7ff57ce6b5	Backend: standardize search values, filters, and sorting for archived items (#1039 ) - all-crawls list endpoint filters now conform to 'Standardize list controls for archived items #1025' and URL decode values before passing them in - Uploads list endpoint now includes all all-crawls filters relevant to uploads - An all-crawls/search-values endpoint is added to support searching across all archived item types - Crawl configuration names are now copied to the crawl when the crawl is created, and crawl names and descriptions are now editable via the backend API (note: this will require frontend changes as well to make them editable via the UI) - Migration added to copy existing config names for active configs into their associated crawls. This migration has been tested in a local deployment - New statuses generate-wacz, uploading-wacz, and pending-wait are added when relevant to tests to ensure that they pass - Tests coverage added for all new all-crawls endpoints, filters, and sort values	2023-08-04 09:56:52 -07:00
Anish Lakhwara	9236a07800	fix: run `yarn format` in frontend dir (#1043 )	2023-08-03 19:12:48 -07:00
Ilya Kreymer	362afa47bd	Support for Public / Shareable Collections (#1038 ) * collections: support toggling collections public/private, viewable via RWP - backend: add 'public' to collection model, support patching to update - backend: add .../collections/<id>/public/replay.json for public access - backend: add CORS handling for public endpoint - frontend: support 'make shareable / make private' dropdown actions on collection detail + collection list views - frontend: show shareable / private icons by collection name on detail + list views - frontend: link to replayweb.page for standalone browsing - frontend: add embed code popup when a collection is shareable - refer to public collections as 'shareable' for now --------- Co-authored-by: Henry Wilkinson <henry@wilkinson.graphics>	2023-08-03 19:11:01 -07:00
sua yoo	62d3399223	Add info bar to Collection detail view (#1036 ) - Adds Collection info bar to detail view - Update "Web Captures" -> "Archived Items" - Updates Collection list columns to match - Refactors `btrix-desc-list` and usage in `workflow-details` to reuse horizontal info bar component	2023-08-03 16:58:56 -07:00
Anish Lakhwara	af09d56ef6	Merge pull request #1035 from webrecorder/backend-init feat: Display waiting message while backend is initializing	2023-08-02 17:39:47 -07:00
Anish Lakhwara	fa58e77167	fix: remove strange character?	2023-08-02 17:34:09 -07:00
Anish Lakhwara	5ed2faaecc	fix: need to use `window.timeOut` to get a timerId back	2023-08-02 17:31:01 -07:00
Anish Lakhwara	6ecfd8ec24	fix: timerId not timeoutId	2023-08-02 17:28:07 -07:00
Anish Lakhwara	3985cf014e	fix: clear timeout on disconnect callback	2023-08-02 17:26:26 -07:00
Anish Lakhwara	196b26c60e	fix: center text	2023-08-02 17:21:36 -07:00
Anish Lakhwara	f1d91e3bf9	fix: add styling	2023-08-02 17:18:40 -07:00
Anish Lakhwara	a8bedeffb5	fix: take Sua's suggestons, less code needed	2023-08-02 17:10:45 -07:00
Anish Lakhwara	2f26fcefce	fix: make pretty & work correctly	2023-08-02 16:36:28 -07:00
Anish Lakhwara	06918c967b	feat: use html dialog instead	2023-08-02 11:37:55 -07:00
Anish Lakhwara	84a60b54e4	feat: Display waiting message while backend is initializing	2023-08-01 17:18:05 -07:00
Ilya Kreymer	45eaa0b3a3	version: bump to 1.6.0-beta.3	2023-08-01 09:48:17 -07:00
sua yoo	cc52dfd940	Sort Collections by size (#1026 ) - Adds "Size" column to Collections list view - Adds "Size" option to sort dropdown	2023-08-01 09:47:47 -07:00
Anish Lakhwara	32428f4d93	fix: usr/bin/env bash interpreter for btrix (#1028 )	2023-08-01 09:28:56 -07:00
sua yoo	54e2b2c703	List web captures in Collection (#1024 ) - Adds tab for "Web Captures" in Collection detail view - Move Collection description under Replay section - Fixes app reloading when clicking into a Collection - Standardizes Web Capture list headers from "Finished -> "Created Date"	2023-08-01 09:14:27 -07:00
Ilya Kreymer	06cf9c7cc3	add crawl ending states: 'generate-wacz', 'uploading-wacz', 'pending-wait' that occur after a crawl is finished or is being stopped (#1022 ) operator: ensure transitions from each of these states is supported, including to 'waiting_capacity' add extra check on stopping to avoid transitioning back to a running state after crawl is finished ui: add states to UI display, localization, add as active states fixes #263	2023-08-01 00:15:59 -07:00
Anish Lakhwara	d848502f84	fix(build): add DOCKER_BUILDKIT=1 to frontend Dockerfile to better support older versions of Docker (#1021 )	2023-08-01 00:15:03 -07:00
Ilya Kreymer	7ea6d76f10	Resource Constraints Cleanup: (fixes #895 ) (#1019 ) * resource constraints: (fixes #895) - for cpu, only set cpu requests - for memory, set mem requests == mem limits - add missing resource constraints for minio and scheduled job - for crawler, set mem and cpu constraints per browser, scale based on browser instances per crawler - add comments in values.yaml for crawler values being multiplied - default values: bump crawler to 650 millicpu per browser instance just in case cleanup: remove unused entries from main backend configmap	2023-08-01 00:11:16 -07:00
Anish Lakhwara	d8502da885	fix(build): use `/usr/bin/env bash` instead of `/bin/bash` (#1020 ) * fix: add to various other shell scripts	2023-07-28 21:50:04 -07:00
Ilya Kreymer	c76dd10928	chart: always pull latest crawler image - since default image is pointing to webrecorder/browsertrix-crawler:latest, makes sense to always pull latest (#1018 )	2023-07-27 12:41:41 -07:00
Anish Lakhwara	a347f61973	ci: password check: fix: don't break on ScannerError (#1017 )	2023-07-27 07:19:27 -07:00
Vinzenz Sinapius	5807507f29	Add proxy settings for crawler and profilebrowser (#997 )	2023-07-26 16:11:10 -07:00
sua yoo	7069b33646	Show only running crawls in superadmin view (#1015 ) - Show separate crawls list for admin view, fixes #1010	2023-07-26 15:48:20 -07:00
Ilya Kreymer	6506965d98	Streaming Download for Collections (#1012 ) * support streaming download of collections (part of #927) - WACZ zip created on the fly using stream-zip - add 'Download Collection' option to collection detail and list - after editing collection, return to collection view - tests: add test for streaming download, ensure WACZ files + datapackage present, STORE compression used --------- Co-authored-by: sua yoo <sua@suayoo.com>	2023-07-26 15:42:17 -07:00
Anish Lakhwara	6062042fae	feat: create DO registry if it doesn't exist (#947 ) - if use_do_registry is enabled and registry doesn't exist, create it	2023-07-26 15:41:03 -07:00
Anish Lakhwara	4c1465d94b	feat: ansible DO teardown (#950 ) * feat: ansible DO teardown * fix(DO): idempotency issues in ansible teardown * chore(DO): remove unused code * docs(ansible): mention teardown in the docs * fix: pass ansible-lint * fix: point database backup upload to the correct location in DO space	2023-07-26 15:38:59 -07:00
Anish Lakhwara	b5a9c42df1	feat: add pre-commit to check we don't have real passwords in yml files (#990 ) * feat: use existing pre-commit framework * feat(ci): add github action for password_check * feat: add some simple tests to password_check.py * fix: set `backend_password_secret` in default values.yaml to an allowed password	2023-07-26 13:29:37 -07:00
Tessa Walsh	c21153255a	Rename notes to description in frontend and backend (#1011 ) - Rename crawl notes to description - Add migration renaming notes -> description - Stop inheriting workflow description in crawl - Update frontend to replace crawl/upload notes with description - Remove setting of config description from crawl list - Adjust tests for changes	2023-07-26 13:00:04 -07:00
Ilya Kreymer	4bea7565bc	load handling: scale up redis only when crawler pods running (#1009 ) Operator: Modified init behavior to only load redis when at least one crawler pod available: - waits for at least one crawler pod to be available before starting redis pod, to avoid situation where many crawler pods are in pending mode, but redis pods are still running. - redis statefulset starts at scale of 0 - once crawler pod becomes available, redis sts is scaled to 1 (via `initRedis==true` status) - crawl remains in 'starting' or 'waiting_capacity' state until pod becomes available without redis pod running - set to 'running' state only after redis and at least one crawler pod is available - if no crawler pods available after running, or, if stuck in starting for >60 seconds, switch to 'waiting_capacity' state - when switching to 'waiting_capacity', also scale down redis to 0, wait for crawler pod to become available, only then scale up redis to 1, and get back to 'running' other tweaks: - add new status field 'initRedis', default to false, not displayed - crawler pod: consider 'ContainerCreating' state as available, as container will not be blocked by resource limits - add a resync after 3 seconds when waiting for crawler pod or redis pod to become available, configurable via 'operator_fast_resync_secs' - set_state: if not updating state, ensure state reflects actual value in db	2023-07-26 08:40:05 -07:00
Tessa Walsh	608a744aaf	Add migration to replace None with 0 for configmap CRAWL_TIMEOUT (#1008 )	2023-07-24 15:49:26 -04:00
Tessa Walsh	fcd48b1831	Add totalSize to collections and make it sortable in list endpoint (#1001 ) * Precompute collection.totalSize and make sortable * Add migration to recompute collection data with totalSize	2023-07-24 13:12:23 -04:00
sua yoo	75b011f951	Upload WACZ via UI (#992 ) - Users can now upload .WACZ archives from the "Archived Data" page. - Can specify name, description, tags and collection(s) to add upload to - Show progress of upload - Support canceling upload	2023-07-21 16:45:52 +02:00
Tessa Walsh	9f32aa697b	Add collections and tags to upload API endpoints (#993 ) * Add collections and tags to uploads * Fix order of deletion check test * Re-add tags to UploadedCrawl model after rebase * Fix Users model heading	2023-07-21 16:44:56 +02:00
Tessa Walsh	4014d98243	Move pydantic models to separate module + refactor crawl response endpoints to be consistent (#983 ) * Move all pydantic models to models.py to avoid circular dependencies * Include automated crawl details in all-crawls GET endpoints - ensure /all-crawls endpoint resolves names / firstSeed data same as /crawls endpoint for crawls to ensure consistent frontend display. fields added in get and list all-crawl endpoints for automated crawls only: - cid - name - description - firstSeed - seedCount - profileName * Add automated crawl fields to list all-crawls test * Uncomment mongo readinessProbe * cleanup CrawlOutWithResources: - remove 'files' from output model, only resources should be returned - add _files_to_resources() to simplify computing presigned 'resources' from raw 'files' - update upload tests to be more consistent, 'files' never present, 'errors' always none --------- Co-authored-by: Ilya Kreymer <ikreymer@gmail.com>	2023-07-20 13:05:33 +02:00
Tessa Walsh	577416024b	Fix pull_request syntax in ansible lint GH Action (#995 ) * Fix pull_request syntax in ansible lint GH Action * Only lint Digital Ocean playbook for now * fix: pass ansible lint --------- Co-authored-by: Anish Lakhwara <anish+git@lakhwara.com>	2023-07-20 12:13:52 +02:00
sua yoo	85913112a2	Upgrade lit + shoelace to reduce build size (#938 ) * upgrade lit * upgrade shoelace * upgrade testing libraries * add webpack bundle analyzer * revert shoelace changes * remove bundle analyzer * remove console log	2023-07-20 11:50:05 +02:00
Tessa Walsh	d5c3a8519f	Add crawler Use Sitemap option to Browsertrix Cloud (#978 ) * Add user-guide docs for Use Sitemap option --------- Co-authored-by: Henry Wilkinson <henry@wilkinson.graphics>	2023-07-19 13:57:52 -04:00
Anish Lakhwara	db851b8360	Merge pull request #976 from webrecorder/ansible-lint-action feat: ansible lint github action	2023-07-19 03:22:22 +10:00
Ilya Kreymer	a5312709bb	fix issues that caused cronjob container to crash: (#987 ) - don't set CRAWL_TIMEOUT to "None" in configmap, and if encountered, just set to 0 - run register_exit_handler() after run loop has been inited	2023-07-18 18:08:53 +02:00
sua yoo	c5b3be0680	Fix frontend formatting pre-commit (#991 ) * update lint staged config * remove prettier defaults	2023-07-18 17:51:13 +02:00
Anish Lakhwara	4fed3ed1b0	fix: resolve ansible pipenv dependencies successfully (#977 )	2023-07-18 17:39:38 +02:00
Ilya Kreymer	5dede47874	remove accidentally added values file!	2023-07-16 15:05:08 +02:00
Anish Lakhwara	bc82f562dc	feat: ansible lint github action	2023-07-10 17:58:47 -07:00
Ilya Kreymer	2372f43c2c	frontend: fix to collection editor with crawls and uploads (#971 ) * frontend: - follow up to #969, fixes crawl workflows by using crawl-specific endpoint and merging results * get crawls and uploads concurrently --------- Co-authored-by: sua yoo <sua@suayoo.com>	2023-07-10 19:29:19 +02:00
sua yoo	f3660839bf	Allow users to add uploads to collections (#968 ) * show uploads in 'Select Uploads' section	2023-07-09 22:21:50 -07:00
Ilya Kreymer	7d694754c6	uploads api ext: (#970 ) - also support collectionId filter on /all-crawls - update tests	2023-07-09 22:12:54 -07:00

1 2 3 4 5 ...

730 Commits