sc_crawler.utils
Functions:
- jsoned_hash – Hash the JSON-dump of all positional and keyword arguments.
- sc_json_serializer – JSON serializer for SQLAlchemy engines writing SC Crawler JSON columns.
- create_sc_engine – Create a SQLAlchemy engine that can persist typed JSON model fields.
- hash_database – Hash the content of a database.
- chunk_list – Split a list into chunks of a specified size.
- scmodels_to_dict – Creates a dict indexed by key(s) of the ScModels of the list.
- is_sqlite – Checks if a SQLModel session is binded to a SQLite database.
- is_postgresql – Checks if a SQLModel session is binded to a PostgreSQL-like database.
- float_inf_to_str – Transform to string if a float is inf.
- table_name_to_model – Return the ScModel schema for a table name.
- get_row_by_pk – Get a row from a table definition by primary keys.
- nesteddefaultdict – Recursive defaultdict.
- list_search – Search for a dict in a list with the given key/value pair.
- convert_gb_to_mib – Convert gigabytes to mebibytes.
func jsoned_hash
jsoned_hash(*args, **kwargs)
Hash the JSON-dump of all positional and keyword arguments.
Examples:
>>> jsoned_hash(42)
'0211c62419aece235ba19582d3cf7fd8e25f837c'
>>> jsoned_hash(everything=42)
'8f8a7fcade8cb632b856f46fc64c1725ee387617'
>>> jsoned_hash(42, 42, everything=42)
'f04a77f000d85929b13de04b436c60a1272dfbf5'
func sc_json_serializer
sc_json_serializer(x)
JSON serializer for SQLAlchemy engines writing SC Crawler JSON columns.
func create_sc_engine
create_sc_engine(connection_string, **kwargs)
Create a SQLAlchemy engine that can persist typed JSON model fields.
func hash_database
hash_database(connection_string, level=HashLevels.DATABASE, ignored=['observed_at'], progress=None, exclude_tables=[])
Hash the content of a database.
Parameters:
- connection_string (
str) – SQLAlchemy connection string to connect to the database. - level (
HashLevels) – The level at which to apply hashing. Possible values are 'DATABASE' (default), 'TABLE', or 'ROW'. - ignored (
List[str]) – List of column names to be ignored during hashing. - progress (
Optional[Progress]) – Optional progress bar to track the status of the hashing. - exclude_tables (
List[ScModel]) – Optional list of tables not to be hashed.
Returns:
func chunk_list
chunk_list(items, size)
Split a list into chunks of a specified size.
Examples:
>>> [len(x) for x in chunk_list(range(10), 3)]
[3, 3, 3, 1]
func scmodels_to_dict
scmodels_to_dict(scmodels, keys)
Creates a dict indexed by key(s) of the ScModels of the list.
When multiple keys are provided, each ScModel instance will be stored in the dict with all keys. If a key is a list, then each list element is considered (not recursively, only at first level) as a key. Conflict of keys is not checked.
Parameters:
- scmodels (
List[ScModel]) – list of ScModel instances - keys (
List[str]) – a list of strings referring to ScModel fields to be used as keys
Examples:
>>> from sc_crawler.vendors import aws
>>> scmodels_to_dict([aws], keys=["vendor_id", "name"])
{'aws': Vendor...
func is_sqlite
is_sqlite(session)
Checks if a SQLModel session is binded to a SQLite database.
func is_postgresql
is_postgresql(session)
Checks if a SQLModel session is binded to a PostgreSQL-like database.
Dialect name is checked for PostgreSQL or CockroachDB.
func float_inf_to_str
float_inf_to_str(x)
Transform to string if a float is inf.
func table_name_to_model
table_name_to_model(table_name)
Return the ScModel schema for a table name.
func get_row_by_pk
get_row_by_pk(session, model, pks)
Get a row from a table definition by primary keys.
Parameters:
- session (
Session) – Connection for database connections. - model (
ScModel) – An ScModel schema definition with table reference. - pks (
dict) – Dictionary of all the primary keys for the row,.
Returns:
ScModel– ScModel object read from the database.
func nesteddefaultdict
nesteddefaultdict()
Recursive defaultdict.
Examples:
>>> foo = nesteddefaultdict()
>>> foo["bar"]["baz"] = 43
>>> from json import dumps
>>> dumps(foo)
'{"bar": {"baz": 43}}'
func list_search
list_search(items, key, values)
Search for a dict in a list with the given key/value pair.
When multiple values are provided, it will use the first field with a matching name with either keys.
func convert_gb_to_mib
convert_gb_to_mib(gb)
Convert gigabytes to mebibytes.
Parameters:
- gb (
int) – Size in gigabytes (GB, decimal: 10^9 bytes).
Returns:
int– Size in mebibytes (MiB, binary: 2^20 bytes).